COSMICS · CHAPTER 6.8
The Planet Has Data Deserts
Observation density follows infrastructure and money, so some regions enter global models with much thinner evidence than others.
EARTHVISION LAB · ~15 MIN READ
GEOGRAPHY OF EVIDENCE
A global dataset can be global in extent and profoundly local in evidence. Weather stations cluster where agencies can build and maintain them. Biodiversity records accumulate near roads, universities and funded field programs. Ocean observations thin with depth and remoteness. Ground validation follows places researchers can reach. The final map can cover Earth evenly while the observations underneath it do nothing of the sort.
This matters because missing data is rarely random. Mountain regions, small islands, conflict zones, deserts, tropical forests and low-income countries can be difficult or expensive to instrument. Those same places may contain climatic and ecological conditions poorly represented elsewhere. A model trained or initialized from an uneven observing network therefore inherits geography before it makes its first prediction.
The phrase data desert sounds as though nature created it. Usually institutions did. Instruments require procurement, power, calibration, communications, technicians, spare parts, standards and years of operating budgets. The sensor may be scientifically mature. The network is an economic system.
GBON
Weather forecasting has an explicit minimum observing network
The World Meteorological Organization created the Global Basic Observing Network, GBON, to define a minimum set of surface and upper-air observations that countries should collect and exchange internationally. The logic is straightforward: numerical weather prediction is global, so one country's missing observations degrade information used beyond its borders.
The network exposes how large the gap remains. In July 2026, WMO reported that across 77 Least Developed Countries and Small Island Developing States, nearly 90% of the observations required by GBON were still missing. Some stations do not exist. Others exist but are non-operational or report inconsistently.
Africa is particularly sparse in conventional surface observations relative to global standards. Small island states face a different geometry: a few stations have to represent large ocean areas while equipment is exposed to salt, storms, expensive logistics and limited technical capacity. A weather balloon remains a simple idea attached to a supply chain.
WMO's Systematic Observations Financing Facility was created specifically because buying an instrument once does not solve the problem. The objective is sustained collection and international exchange. An observation network that works for eighteen months and then loses maintenance funding is not climate infrastructure. It is a short experiment.

GLOBAL CONSEQUENCES
A missing station damages forecasts far from the station
Weather provides a rare opportunity to quantify the value of filling a data desert because forecast centres can run observing-system experiments. Remove or add classes of observations, rerun the assimilation and compare the resulting forecasts. The value of a station is measured by how much it changes the global estimate of atmospheric state.
WMO cites an ECMWF impact experiment estimating that closing essential observation gaps could reduce forecast errors by about 30% in Africa and 20% in the Pacific. The improvement is not confined to the country operating the station. Atmospheric information moves with the weather and enters global data-assimilation systems used by forecast centres everywhere.
This is one reason observations are exchanged internationally under meteorological agreements. A radiosonde launched in Bhutan or Madagascar can improve the initial state of a global model whose forecast later matters elsewhere. The instrument has a location. Its informational value does not respect the same border.
Satellite observations reduce many geographic gaps, especially over oceans and remote land, but they do not make the in-situ network redundant. Surface pressure, upper-air profiles, precipitation gauges and reference measurements constrain and calibrate variables that satellites observe differently or indirectly. Global forecasting is a hybrid system because neither layer is sufficient alone.
TAXONOMIC AND GEOGRAPHIC GAPS
Biodiversity has darkspots rather than empty continents
Biological data deserts are harder to define because there is no equivalent of one globally agreed weather-station spacing. A region can be well sampled for birds and poorly sampled for fungi, well mapped for forest cover and poorly inventoried for understory plants. The gap has both geography and taxonomy.
Kew's 2023 State of the World's Plants and Fungi work identified 32 plant-diversity darkspots where major knowledge gaps coincide with high expected botanical diversity. The concept is useful because it does not ask where no data exists. It asks where the mismatch between expected importance and available knowledge is unusually large.
Citizen-science platforms can add enormous numbers of records, but participation follows people, roads, connectivity and interests. Museum collections inherit centuries of expedition history. Molecular surveys require laboratories and reference databases. A record density map therefore contains the history of science as well as the distribution of species.
This is why absence in a biodiversity database is particularly dangerous to interpret literally. A species may be absent, undetected, undescribed, unsequenced, geographically inaccessible or simply uninteresting to the people who have sampled the area so far. The database has a blank cell. Biology offers several reasons.
THE MISSING LAYER
Observation coverage is infrastructure and should be treated like it
Data deserts persist when observing systems are treated as projects rather than infrastructure. A research grant can deploy sensors. Long-term records require maintenance, communications, quality control, metadata, calibration and institutional ownership after the grant ends. The expensive part is often not discovering how to measure the variable. It is continuing to measure it for decades.
The same issue appears in ground truth for satellite and AI systems. A model may generate predictions everywhere, but validation sites can remain concentrated in a few countries and land-cover types. Global deployment then expands faster than global evidence. The output surface grows. The independent test network does not necessarily follow.
A mature planetary system should therefore treat observation density, reporting reliability and validation coverage as first-class infrastructure layers. Where are stations operating? Which variables do they measure? When did each last report? Where does a product rely mostly on interpolation? Which ecosystems have independent validation? Those questions belong beside the environmental values, not in a technical appendix nobody opens.
A global model can calculate everywhere. The observing network that keeps it honest still has addresses.
