← COVER

The Sensors Stopped Speaking Separate Languages

DRISHTI · CHAPTER 5.2

The Sensors Stopped Speaking Separate Languages

Radar, optical light, weather, elevation and a field report can now enter one model. The hard part is remembering afterwards which of them actually saw what.

EARTHVISION LAB · ~17 MIN READ

Take one agricultural field. An optical satellite sees reflected light. Radar responds to structure and moisture. A weather station measures the air a few kilometres away. A soil probe samples one point every few minutes. A yield monitor on the combine records what happened at harvest. A farmer's note says the east corner flooded in May. All of them can be correct and still disagree.

For decades these were processed in separate pipelines, and not because of file formats. They describe different physical properties in different mathematical spaces. A red band and a near-infrared band come from the same family of instrument. Radar and optical sensors have different imaging physics altogether. Elevation hardly changes, weather changes hourly, and text refers to objects and causes rather than measurements.

Two related problems follow. Multimodal learning asks whether such different data can be encoded together in one shared space where they inform one another. Fusion asks the operational version: given several measurements arriving at different times with different errors, what should the system believe about this field now? The field has one physical history. The instruments have several partial versions of it.

Illustration of a combine harvester's yield monitor.
View: A yield monitor records what happened during harvest.

The hard part is deciding what corresponds to what

Multimodal models usually begin with a separate encoder for each kind of data, each turning its input into an internal representation. Training then encourages representations of the same place, time or object to become mutually informative. The difficult word in that sentence is same.

Space rarely lines up. Radar and optical images differ in geometry, viewing angle and resolution. A weather value may describe a grid cell covering thousands of optical pixels, and a probe a few cubic centimetres of soil. Resampling can put them all on one grid. It cannot make them describe the same volume of Earth, and a coarse value copied into thousands of fine pixels has acquired detail in the file and none in reality. Time is worse: a storm can change the radar signal between two optical passes, and a field note written in the morning may describe a different field from the one the satellite saw that afternoon. Treat them as simultaneous and the model learns disagreements created by the clock rather than by the Earth.

Meaning is hardest of all. A note saying that flooding blocked the eastern road links an event, an object and a direction. An image contains no word road. The model has to learn that a line in one representation corresponds to a named thing in another, at which point multimodal learning starts to look less like data fusion and more like translation.

Where the sources meet inside the model matters too. Combining them early captures fine relationships but assumes every input is present. Combining them late keeps each source's structure and copes with gaps, at the cost of compressing some information before the sources ever interact. Cross-attention lets one source query another for what it needs: optical data can consult radar where cloud makes it unsure, and text can attend to the image region it names. The choice determines whether a model treats disagreement as noise, keeps it as evidence or routes around a missing input, and a planetary system will meet all three before lunch.

Illustration of a field reference observation.
View: A field observation may not align with the satellite's footprint or its time of observation.

TerraMind made generation part of the reasoning

ESA and IBM released TerraMind in April 2025 as a multimodal Earth-observation foundation model. The research paper describes pretraining across nine geospatial modalities from a large, globally distributed dataset, and ESA's summary describes more than nine million samples spanning radar, optical imagery, topography, land cover, elevation, location and related data.

TerraMind works at two levels at once. A token-level representation carries broad context efficiently, while a pixel-level path keeps the fine spatial detail that matters for outlining fields or flood edges. The model does not have to choose between seeing the whole region and seeing the boundary before it starts learning.

Its stranger idea is called Thinking-in-Modalities. The model can generate a modality it was not given, a radar-like view or a land-cover map, and use that generated layer as an intermediate step while solving a task. A missing view becomes a temporary hypothesis that helps it reason across the sources it does have.

That is technically interesting and epistemically dangerous in exactly the right way. A generated modality is not an observation. It is the model's estimate of what an observation would probably have shown, useful because the learned relationships are strong. If the rest of the system forgets that distinction, synthetic evidence can quietly acquire the authority of measured evidence, and nobody will be able to say afterwards which instrument it came from, because none did.

Illustration of different kinds of evidence about one place.
View: The same place can arrive as optical imagery, radar, rainfall, temperature, labels and field reports.

Fusion is mostly a problem of weighting evidence

Fusion is not averaging the answers. Optical greenness, radar backscatter, heat emitted by the surface and microwave brightness are not four noisy readings of one variable. They respond to different properties, and the relationships between them shift with vegetation density, surface roughness, soil texture, viewing angle, temperature and water content. Errors can also be shared: two satellite products may lean on the same atmospheric correction or the same weather dataset, and treating them as independent witnesses makes the combined answer more confident without making it more correct. A crowd of witnesses is less impressive when they all heard the same rumour.

State estimation gives fusion a more explicit logic. The system carries a prior estimate of the current state; new observations arrive; the estimate moves towards each according to how uncertain the prior and the observations are believed to be. Volume 4's data assimilation is the mature physical example. The hard part is never the update equation. It is the error model. Assume an instrument is more precise than it is and it drags the state; underestimate model error and observations barely move it; treat correlated errors as independent and the same evidence gets counted twice. Learned fusion can infer the weighting from data instead, and can also learn the laziest shortcut, ignoring every source but its favourite until deployment reaches the day its favourite fails.

NASA's SMAP Level-4 soil-moisture product shows fusion working at planetary scale, and the result looks deceptively simple: surface and root-zone soil moisture on a 9-kilometre global grid every three hours. Volume 3 noted that the SMAP satellite reads faint microwaves from the top few centimetres of soil. The Level-4 system assimilates those observations into NASA's Catchment land-surface model, which carries water through time and down through the soil, while the satellite corrects the state wherever it has looked.

The root-zone value is fusion in the strongest sense: what a radiometer can observe near the surface, combined with rainfall, land-surface physics and the model's account of water moving downward. The user receives one field, but different parts of it owe different shares of their value to measurement and to model. NASA's Version 8 continues refining the rainfall inputs, land parameters and the microwave model, and the product improves because both witnesses improve: the instrument and the model used to interpret it.

A good fusion system preserves the argument

Several sensors are valuable partly because they fail differently. Cloud can ruin an optical image while radar carries on. Radio interference can spoil a microwave measurement while a ground station is fine. A probe can fail locally while a satellite still sees the wider pattern. Agreement between genuinely independent sources raises confidence. Disagreement should not automatically be averaged away: if radar says the surface changed and optical imagery says it did not, the mismatch may point to moisture, geometry, vegetation or simply a timing difference, and it can carry more diagnostic information than either answer.

Compression can erase the reason the views differed. Radar may be noisy because of speckle, optical data may be cloud-affected, a field report may be late or subjective, and a generated modality may never have been measured at all. For a low-stakes map, the combined answer may be enough. For a system that will trigger another observation, change a forecast or influence a decision, the lineage has to survive: which source contributed, when it was observed, whether it was measured, inferred or generated, and how much the answer changes when that source is removed.

This is the difference between a shared representation and a shared truth. The representation can be excellent while one contributor is stale or wrong, and a system that cannot recover the lineage becomes confident about evidence whose origin it cannot inspect. Provenance therefore belongs inside the state, not in a log file nobody reads.

Once a system can hold all these sources together, two larger things become possible. It can reuse what it learned about the planet for many tasks at once, and it can search hundreds of variables for relationships nobody thought to look for. Before either, it had to prove itself on the oldest prediction problem there is.

Fusion should reduce uncertainty without erasing the evidence that created it.
Illustration of a failed soil probe in a field.
View: A probe can fail locally while a satellite still sees the wider pattern.