2.1 How spatial evidence is generated

In 1.3, we looked at how spatial evidence is produced and how different data sources, methods, and assumptions can all affect what a map is showing and how confidently it can be used.

This section builds on that foundation in three parts: a general workflow of raw data or observations to evidence, different formats of spatial evidence, Earth observation as an increasingly important source behind many spatial evidence products, and, finally, the specific steps used to generate the datasets behind this hub, from field to open data.

Figure 1: Spatial observations from the field can be coupled with satellite and other datasets to create informative, rich maps.

2.1.1 From observations to evidence

In the context of landscape restoration, spatial evidence is the analysed or modelled output derived from spatial data (data with locational information) that helps describe landscape conditions and trends, and inform decisions. It may be presented in a range of formats including as static or interactive maps, charts, plots, in dashboards, summary tables, and other visualisations. This hub focuses mainly on spatial evidence presented as maps, though the same principles apply to interpreting any of these formats.

Spatial evidence is ultimately a product of different kinds of observations of the world coming from a wide variety of sources - satellites, aircraft, weather stations, direct field samples, IoT sensors, household surveys, citizen-science data, and more.

Getting from a raw observation to a usable piece of evidence almost always involves some degree of processing, transformation, or combination with other data. In our data-rich world, spatial evidence products increasingly combine several sources to produce richer, more informative outputs.

Observations to spatial evidence

Observe Collect data/observations of the world
satellites · field data · sensors

Process Prepare and quality-control data
correction · cleaning · calibration

Analyse / model Transform, combine, analyse data
indices · summaries · comparisons · predictions

Product Represent the information across space
maps · indicators · datasets

Evidence Interpret in relation to a question or decision

Spatial evidence as raster (and other) data

Spatial evidence is usually represented in one of two main formats: raster or vector.

A raster represents a landscape as a continuous and regular grid of cells (pixels), with each cell storing a value. This makes raster data the best format for capturing continuous and sometimes categorical variables.

A vector represents discrete features using either points, lines and polygons. Points might represent field observations or sampling sites; lines could represent rivers or roads; and polygons might represent fields, protected areas, administrative boundaries or restoration project areas.

Figure 1. Raster data represents a landscape as a grid of cells; vector data represents discrete features as points, lines, and polygons.

Why do we work so much with raster data? Environmental conditions are typically continuous variables - meaning their observable properties tend to change and vary continuously across whole landscapes. A raster allows us to represent this continuity by storing a value for every pixel across the landscape.

This is also how much of our spatial evidence is generated - Earth observation satellites measure the Earth’s surface as grids of pixels and, often as a function of that, statistical and machine-learning models used to estimate environmental conditions across areas often produce their results in the same gridded format. Raster data is particularly useful for comparing conditions across large areas, identifying spatial patterns and trends, and combining multiple environmental indicators.

Raster and vector data are often used together. For example, we might use a raster map to visualise vegetation condition across a whole country and overlay vector boundaries to summarise values within districts, restoration sites or other areas of interest. Read more about spatial data data formats here.

Earth observation: a growing, increasingly essential input

Earth observation (EO) - repeated satellite and aerial imagery of the Earth’s surface - is one of the most important data sources feeding into spatial evidence. Historically, it was expensive and controlled by a small number of national space agencies but over the past couple of decades this has changed substantially with a rapidly growing number of public and commercial satellite missions, combined with free and open data policies from major providers, meaning that consistent, repeated observations of the Earth’s surface are now openly available, at increasing spatial and temporal resolution, across almost the entire globe.

EO data on it’s own can sense and measure physical properties of the Earth’s surface like reflected light, radar backscatter, but is incapable of measuring things like soil organic carbon, woody biomass, or biodiversity, all of which typically require ground-sampling to train and validate models. The combination of EO and field data is very powerful, enabling us to produce spatial evidence that is both spatially continuous and grounded in real-world measurements.

Different degrees of transformation

Direct measurements Direct measurements Satellite reflectance, rainfall measurements, GPS points and field sample locations are examples of data that typically undergo minimal transformation although they may already have undergone important steps such as calibration, correction and quality control. A pixel on a map in such case may be a rainfall measurement (often interpolated from a network of weather stations) or a satellite reflectance measurement at a particular wavelength, for example. Derived indicators Derived indicators Other variables can be calculated from these basic measurements. Vegetation indices such as EVI or NDVI, for example, combine reflectance values from different wavelengths to provide an indicator of vegetation greenness. Temporal summaries & trends Temporal summaries & trends A dataset can also be summarised through time and across space. Raw precipitation measurements can for example be used to compute annual average rainfall, or a statistical trend for each pixel to identify areas where vegetation appears to be increasing or declining. Modelled or predicted variables Modelled or predicted variables Often a variable can't be observed or remotely-sensed continuously across the landscape at all. In such cases, field observations can be combined with environmental variables from satellites using statistical or machine-learning methods to estimate conditions in places where no direct measurement exists. Pixels are assigned a predicted value, and often an associated uncertainty, based on the model's understanding of how the variable behaves across the landscape.

From direct measurements to derived indicators, temporal summaries, and modelled predictions, progressively more processing steps are involved and, with each step, more assumptions are introduced and more opportunity exists for error to propagate and compound. 2.2 picks this up in more detail, working through concrete examples and technical considerations to make when using spatial evidence to support decisions.

2.1.2 The K4GGWA evidence pipeline

The general process described above applies directly to how the evidence products available on the K4GGWA platform are generated. Below we step through that pipeline, from systematic field sampling through laboratory analysis and modelling, to open data access and decision-making.

1. LDSF - systematic field sampling

Field data collection follows the Land Degradation Surveillance Framework (LDSF) - a systematic, nested sampling design used across Africa's diverse landscapes, including throughout the GGW countries.

Sites are selected at random within a landscape. Each site is a 10km x 10km quadrat containing 16 clusters, each made up of 10 plots, each in turn divided into 4 subplots. This nested, multiscalar design captures spatial variation at several scales within the same landscape, which keeps the resulting data minimally biased and makes it a robust foundation for the modelling described later in this pipeline.

The framework is designed to collect soil and land health data like erosion prevalence, soil physico-chemical properties, and biodiversity, but may be extended with additional modules depending on the context, for example, using the rangeland or eDNA modules, so the same underlying design can be adapted to different landscape and monitoring needs.

2. LIMS - laboratory analysis and spectroscopy

Alongside physical soil samples, field teams record observational data directly through a digital (ODK) form including erosion prevalence and other visually assessed indicators, identified plant species, biomass measurements, and more. This means every physical sample is paired with rich, geolocated field observations.

Physical samples are transported either to the main laboratory facility at World Agroforestry in Nairobi, or to one of a growing network of regional soil labs part of ongoing efforts towards regional lab capacity building.

In the lab, samples are processed through our Lab Information Systems (LIMS), and analysed using state-of-the-art spectroscopy alongside other equipment, enabling high-throughput, high-accuracy analysis at a scale that would not be feasible using conventional wet-chemistry methods. Read more here.

3. Building the LDSF database

Laboratory results and field-collected (ODK) observations are processed, cleaned, and quality-controlled, then integrated into the LDSF database.

This database, which comprises Africa's largest soil spectral library, alongside associated plant and land health data, are the reservoir that all downstream analytical products are drawn from. Everything described in the next two steps depends on the quality and consistency established at this stage.

4. Analyses and predictive modelling

The LDSF database is used to train deep learning models of land health combining field-referenced measurements with Earth observation imagery and other environmental covariates (see 2.1.2) to predict conditions continuously across the landscape, not just at sampled locations.

This is the same points-to-predictions step introduced in the general workflow above, applied at scale, using advanced models to generate continuous raster products together with explicit measures of uncertainty, for visibility over predicted value at any location and confidence in the output.

5. Open data access

The resulting data products are published through an open, STAC-based data infrastructure, including a dedicated GGW catalog that is interoperable, transparent, and FAIR, such that maps, metadata, and underlying assets can be reused across different workflows and decision contexts.

For many GGW-relevant indicators, these products outperform available global default datasets, and technically capable practitioners can go further still by combining and building on the underlying data to generate their own evidence.

I the next section 2.2 we explore practical questions and considerations when integrating evidence into our decisions - covering the quality and possible biases of the underlying data, transformation steps involved in producing the final output, metadata availabilities to verify any of this, and practical questions about the evidence itself including what variable is represented, the unit of measurement, the spatial and temporal resolution, and how it might be combined with other evidence to support a more informed decision.

Learn more

Back to top