DRISHTI · CHAPTER 5.1
When the Rules Became Weights
We stopped telling machines what a field looks like and showed them examples instead. Then they learned that a field is something that happens in a particular order.
EARTHVISION LAB · ~16 MIN READ
THE OLD ARRANGEMENT
For a long time, automated Earth analysis meant turning an expert's judgement into instructions. Vegetation could be separated from bare ground with a ratio of two bands. Water could be isolated with a threshold. Texture could tell a city from a bright desert. The computer did the calculation very quickly, but the idea of what mattered still came from a person.
That arrangement was more powerful than it now sounds. An index or a decision rule can be inspected, carried to another continent and usually explained in physical terms. If the output looks wrong, an analyst can ask which threshold fired and why. The method becomes awkward only when the useful signal is spread across dozens of bands, neighbouring pixels, seasons and contexts that refuse to reduce to one rule.
Machine learning changed the division of labour. Instead of specifying the pattern, people supplied examples and a goal, and the model adjusted its internal parameters until the examples separated. The rules did not disappear. They became millions, and then billions, of weights that no committee was ever expected to review. Civilisation accepted this surprisingly quickly.
THE TECHNICAL SHIFT
Feature engineering became feature learning
Machine learning did not begin with deep neural networks. Remote-sensing teams used logistic regression, support-vector machines, random forests and other statistical classifiers for years. These could learn a boundary between classes from labelled examples, but their inputs were still usually designed by people: vegetation indices, band ratios, texture measures, elevation, seasonal statistics.
Deep learning moved the learning one step earlier. A convolutional network learns its own filters, responding first to edges, then shapes and textures, then larger arrangements, directly from imagery. A road is not only a particular colour. It is narrow, connected and continuous across space. A settlement is not one bright pixel but a pattern of roofs, streets, shadows and geometry, and a network can learn that pattern without anyone describing it.
That matters because Earth categories are rarely defined by one measurement. Bright rock can resemble concrete. A dry riverbed can resemble a road. A freshly harvested field can resemble bare ground. Context lets a model use relationships that would be tedious to write as rules and brittle to maintain across continents.
The price is interpretability. In a hand-built system, a feature has a name before the model uses it. In a learned system, a useful internal feature may have no human name at all. It exists because adjusting it reduced the training error. The system gained expressive power and lost the comforting fiction that every useful distinction had already been named.
DYNAMIC WORLD
A global classifier that keeps updating, and where it learned
Dynamic World shows what the shift looks like once it becomes infrastructure. Built by Google and the World Resources Institute, it runs a fully convolutional neural network on every new Sentinel-2 image and produces land-cover predictions at 10 metres in near real time, across nine classes: water, trees, grass, flooded vegetation, crops, shrub and scrub, built area, bare ground, and snow or ice. Training still depended heavily on people. The project assembled about 20,000 hand-labelled image tiles spread across the world's biomes, with annotators using sharper imagery and other references to decide what each pixel was. Learning removed the need to write the rules, not the need to define the categories.
Because it runs whenever a usable image arrives, the product stopped being a map published every few years and became a stream of probabilities tied to individual acquisitions. It also reports a probability for every class rather than only the winner. A place may look mostly like crops while keeping a substantial chance of being grass or bare ground. That ambiguity is not an inconvenience to be hidden. It is information about what the model actually saw, and it becomes essential once systems start passing beliefs to one another.
A learned feature is only as good as the examples that taught it, and this becomes a geography problem almost immediately. Roof materials, field sizes, irrigation patterns, crop calendars, cloud regimes and sensor conditions all differ by region. The Dynamic World paper reports the expected unevenness: stable classes such as water and trees are easier; mixed or transient classes such as scrub, bare ground, crops, grass and flooded vegetation are harder; and missed cloud or shadow can turn into confident land-cover errors. The model is global. Its uncertainty is not evenly spread across the globe.
The WILDS benchmark made the general problem explicit by testing models under real shifts, including satellite images separated by geography and time. Standard models did substantially worse on data unlike their training sets, so a random split between training and test data can produce a reassuring score for a situation nobody will ever deploy into. Geography is not noise around a dataset. It is part of how the dataset was made, and a model that works on average can be systematically weak in exactly the places where labels and institutions were already sparse.

TEMPORAL LEARNING
One image is a state. A sequence has behaviour.
A classifier can learn that a mix of texture, geometry and reflectance usually means cropland without knowing anything about planting, irrigation, soil or markets. For land cover, that is fine. The trouble is that many Earth classes are ambiguous in a single frame. Bare soil can be a construction site, a harvested field or a dry lake bed. Water can be permanent, seasonal or a flood. Green can be a crop, a pasture or regrowth. The answer usually lives in what happened before and after, which means the model has to learn that the order of observations matters.
An early way to use a year of images was to collapse it into statistics: peak greenness, lowest moisture, seasonal range. This is efficient and sometimes excellent, and also destructive, because two sequences with the same peak and trough can describe entirely different events if one changed gradually and the other overnight. Recurrent networks answered by carrying a hidden state from one observation to the next. Temporal convolutional networks look for patterns across neighbouring dates, and Pelletier and colleagues showed that such a network could learn when a signal changed, how fast and for how long, straight from a satellite time series without a hand-made seasonal summary.
Attention made the choice of dates explicit. The Pixel-Set Encoder with Temporal Attention Encoder can give more weight to informative dates and less to ordinary ones. A crop field does not need every day of the season to explain itself equally: planting, rapid growth, peak canopy and harvest carry more information than the weeks between. The model learns that calendar without ever being handed one.
Earth does not supply evenly spaced frames. Optical images vanish behind cloud, satellites pass at different intervals, stations fail, and a field survey happens twice a season while a soil probe reports every fifteen minutes. Carrying the last observation forward through a three-week gap silently assumes nothing changed. Interpolating assumes change was smooth. Modern sequence models encode observation times, gaps and masks of what is missing, which does not solve missing data but stops the pipeline pretending it never happened. The gaps carry information of their own: monsoon cloud is seasonal, smoke hides fires, and a model that treats every missing scene as random can erase part of the process it was meant to learn.

CHANGE AND DRIFT
Memory is useful until history changes its rules
Time changes the question from what is here to what is happening here. That sounds like a small grammatical improvement and is actually a different class of problem. A reservoir that is low but rising means something different from one at the same level and still falling. A building site bare for six months means something different from one cleared yesterday. Classic change detection compares two dates. A sequence model can estimate when a change began, whether it lasted, whether it reversed and whether it fits the place's normal seasonal cycle, so a temporary flood, a new lake and a cloud shadow stop looking like the same dark pixels.
This is the difference between mapping and monitoring. A map assigns a state. A monitoring system maintains a state and updates it as evidence arrives. A static map can be excellent and still be out of date. A temporal model is expected to notice when its own previous answer has expired.
That expectation brings a harder problem with it: concept drift. The patterns used to recognise a phase can themselves change. Planting dates shift, reservoirs are operated differently, cities spread and climate moves the seasons. A model trained on yesterday's normal can eventually label tomorrow's normal an anomaly with great confidence. Earth makes this unusually awkward because the system is being altered while the model watches.
A useful planetary system therefore needs two kinds of memory: enough history to know what usually happens, and enough scepticism to notice when the usual relationship has stopped holding. The second cannot come from time alone. It needs other kinds of evidence, arriving through other physics, that can disagree with the sequence in informative ways.
The first machine-learned Earth was a collection of patterns. The next one had to remember which pattern came before which, and then doubt it.
