COSMICS · NOTE 001.5
The Species Nobody Had Time to Count
Camera traps and field recorders now produce more evidence of wildlife than any team could ever review by hand.
EARTHVISION LAB · ~14 MIN READ
THE REAL BOTTLENECK
A motion-triggered camera in a forest does not know what it is photographing. It only knows something moved, so it fires, whether that something is a jaguar or a branch swaying in the wind.
Multiply that camera by the thousands now deployed across protected areas, running for months between visits, and the bottleneck stops being detection. It becomes review. A single research network can accumulate millions of images in one season, the large majority of them blank triggers: wind, shifting light, an animal that had already left the frame before the shutter caught it.
Ecologists have always accepted that trade. A camera never tires and never scares an animal off by being there, but every frame it produces still has to pass in front of a trained eye before it becomes usable data. For decades, that eye, not the camera, was the actual limit on how much of a landscape could be watched at all. The same limit shows up in a different form with audio recorders left running in the field, and the shift worth paying attention to is not that a machine can suddenly notice something a person could not. It is that the machine can get through all of it, which nobody could previously afford to do.

WHAT A CAMERA SEES
The photograph was never the hard part
Google's SpeciesNet is an image classifier, a model trained to sort a photograph into one of a fixed set of categories, built on an architecture called EfficientNet V2 M and trained on more than 65 million camera-trap images pooled from the Wildlife Insights research community and public archives. It sorts into more than 2,000 labels: a species where the image supports one, a broader group such as felidae or mammalia when it does not, and non-animal categories like blank or vehicle for the frames that were never wildlife to begin with.
The honest numbers matter more than the headline. Google reports that the model detects an animal's presence in 99.4% of the photos that actually contain one, reaches an identification at the species level 83% of the time, and gets 94.5% of those species-level calls right. Multiply it through and roughly one photo in five still needs a person to finish the identification. That is not a small gap, but it is a very different problem from the one it replaced: a researcher at Wake Forest University used the model to work through an 11 million photo backlog in a matter of days, and a camera network in Ecuador scaled to 446 cameras and more than 100,000 images in a single year after adopting it, coverage neither project could have reviewed by hand at that pace.
What changed is not that the model replaced the ecologist. It re-routes their attention. The four confident calls out of five get sorted out of the queue, which leaves the finite number of trained reviewers free to spend their time on the contested fifth, and on the rarer species the training data barely covers in the first place. A backlog that used to mean data nobody had gotten to yet now means data a person only has to check, not first discover.
WHAT A FOREST SOUNDS LIKE
Nobody was going to listen to all of it
A passive acoustic recorder left running in a forest or reef for weeks produces audio no team could review at listening speed and keep the rest of their job. Google DeepMind's Perch was trained to classify nearly 15,000 species from sound, mostly birds, along with frogs, crickets, grasshoppers, and some mammals, and to generate embeddings: a compressed numerical fingerprint of a sound clip that other tools can compare against each other, useful for tasks well beyond the species label alone.
Hawaii's honeycreepers, native forest birds under severe pressure from avian malaria and habitat loss, several species already extinct and others down to a few hundred individuals, are exactly the kind of case where review speed decides how much ground a small team can cover. The University of Hawaii's LOHE Lab used Perch to identify honeycreeper calls in field recordings nearly 50 times faster than the conventional method of scanning spectrograms by eye, letting the same handful of researchers monitor more sites than they could ever have covered by ear alone.
The clearest case for using a model at all shows up underwater. Perch's updated version was trained on essentially no underwater audio, yet the same embeddings transfer well enough to flag meaningful events in coral reef soundscapes, because a healthy reef and a degraded one sound different, in the density of snapping shrimp and fish calls, before either condition is visible to a diver. No dive team or hydrophone crew was ever going to log that continuously across a reef system. The recorder already could. Until a model could sort what it captured, that continuous record had no way to become usable data at all.
TELLING ONE ANIMAL FROM ANOTHER
A pattern behind the gills becomes a name
Population science needs more than a count of animals seen. It needs to know whether the whale shark in this month's photo is the same one photographed off the same coast last year, which is the difference between a real population estimate and a number that just counts sightings twice. The nonprofit Wild Me built Wildbook, now branded Sharkbook for sharks, around exactly that problem: the pattern of spots behind a whale shark's gills is as individual as a fingerprint, and the platform matches a new photo against a global library using an approach built the way face-matching software is, only trained on skin patterns instead of faces, returning researchers a ranked list of likely matches rather than a single automatic verdict.
The system extends past the photos researchers submit on purpose. It checks public video uploaded and tagged as whale shark footage, reads the description, runs the same matching model to check for a known individual, and uses location data to filter out anything clearly filmed in an aquarium. A tourist's holiday clip becomes an unplanned entry in a global mark-recapture study, the same kind of repeat-sighting data biologists once relied on physical tags for, without the tourist ever finding out.
None of this replaces the biologist's judgment. A ranked list of candidate matches still needs a person to confirm the top one, and the whole system depends on people continuing to photograph and upload the animals in the first place. What it changes is scale: years of manual photo-matching become weeks, across a population no research team could ever physically travel to in full.

WHAT HASN'T CHANGED
Attention, not understanding
A species label, a matched call, a confirmed individual: all three systems are doing the same job at a scale no person could match. They turn an unmanageable pile of raw evidence, photographs, hours of audio, uploaded video, into a much smaller pile a trained ecologist can actually look at.
None of them explain why a population is falling or a habitat is failing. A species identification is still only a better observation, not an account of the mechanism behind a decline. That harder question, the one this volume keeps returning to, still belongs to the person deciding what the cleared backlog actually means.
A model that finally has time to look at everything still has to hand the interesting cases back to someone who knows what they are looking at.