Often throughout the history of science, new tools advance more rapidly than the science that uses them. I wrote that sentence with four colleagues in 2009, at the end of a paper about the future of environmental sensor networks. I had no idea how literally it would come true.

This week the tool in question is artificial intelligence, and the headlines are breathless. Two mathematicians, working with large language models, produced a proof bearing on the Navier-Stokes equations, one of the seven Millennium Prize problems. A commercial lab then extended the result over a weekend. The blogosphere declared that machines had begun to think like Einstein, and that the rest of science was next.

The same week, Nature published a long feature by Philip Ball asking a sharper version of the question: could an AI trained only on what was known in 1911 rediscover general relativity? Several research teams have now tried variants of this experiment, building so-called historical language models. Their results are sobering. The training data leak: a model supposedly ignorant of everything after 1930 cheerfully answers questions about Franklin Roosevelt's presidency. And when the models do produce physics, they produce it the way a fog machine produces fog, generating plausible sentences with no evident grasp of what they mean. Another team, feeding a model clean synthetic data on planetary orbits, found it never inferred Newton's single law of gravitation. It invented a different wrong law for each solar system.

The Navier-Stokes result and the failed relativity experiments are not in tension. They describe the same machine from two sides. Mathematics is the one domain where every step of an argument can be checked by another program, so a system that generates thousands of candidate ideas and lets a verifier throw out the bad ones can succeed without understanding anything. The mathematicians who got the proof said the English-language explanation the models produced was barely readable. Physics has no such verifier. Ecology, my own field, has nothing remotely like one.

So what are these machines actually good at? The answer arrived the same week, from a neighboring discipline, and it is the reason I have been thinking about campfires.

The campfire

For twenty-six years I directed the James San Jacinto Mountains Reserve, a University of California field station in the mountains east of Los Angeles. In 2002 it became the terrestrial test bed for the Center for Embedded Networked Sensing, a National Science Foundation center co-led by the computer scientist Deborah Estrin. Engineering students, postdocs, and faculty came up the mountain to bury sensors in the soil, hang cameras in nest boxes, and string robots on cables between the pines.

At night we sat around the fire and imagined where it led. Not three hundred sensors on one mountain but tens of thousands across a continent, cheap and self-organizing, feeding a shared system that would learn the pulse of the land the way a physician learns a patient.

The dream had an institutional vehicle. Estrin and I co-chaired the sensors and sensor networks subcommittee for the National Ecological Observatory Network, the largest ecological facility the NSF had ever contemplated. With more than a hundred scientists we argued for the massively distributed version: dense, numerous, adaptive. NEON went another way. By the time my colleagues and I described the plan in a 2007 paper, the design had already settled on clusters of a few instrumented sites per region. What was eventually built, at a cost approaching half a billion dollars, is eighty-one exquisitely calibrated locations across the United States, each with a flux tower, soil arrays, aquatic sensors, standardized field surveys, and annual airborne imagery. Battelle now operates it under a five-year, $416 million award that runs through 2028, a quarter of the way into a thirty-year design life.

Rereading my own papers this week was humbling. The 2007 paper says, in its conclusions, that effective embedded sensing is not always about the largest number of the smallest sensors, and that robust systems need a layered mix of observational resources. The 2009 paper says the biggest limitation is funding, and that prototypes are worthless without mature tools that ordinary scientists can use. We were right on both counts, and the layered system got built. It just got built by different parties who have never spoken to one another.

The discipline that got its NEON

On September 3, Google DeepMind released WeatherNext 3, a machine-learned global weather model that now tops the independent Operational WeatherBench leaderboard, ahead of the European Centre's physics-based ensemble that has been the gold standard for decades. It is, by any measure, the most successful physical-world AI on the planet.

Here is the detail that matters. WeatherNext 3 does not learn from theory. It trains directly on decades of observations from airport weather stations, regional Mesonets, and ships and buoys, alongside raw satellite imagery. It learned the atmosphere from the sensor web that meteorologists built, station by station, over a century. Meteorology had its NEON, dense and continental, long before anyone thought to call it that. The machine simply read it.

This is the inductive engine Ball's article describes. It is not Einstein. It cannot tell you why a front stalls. But given the ground truth, it forecasts better than the equations do, and that is not a small thing.

Ecology never built the sensor web. And yet, while NEON was raising its towers, the distributed layer arrived anyway, through the side door. Tens of millions of people now upload bird sightings to eBird and photographs of plants and insects to iNaturalist. Open weather services interpolate temperature and rainfall onto grids covering every point on Earth. None of it is calibrated; all of it is effort-biased and noisy. But it is continental, continuous, free, and, unlike a federal facility, immune to any single budget line.

What the machine cannot do, yet

One story from the mountain first.

In June 2003 a pair of violet-green swallows nested in box 8 at the James Reserve. Our sensor network recorded that on one day the daytime high fell to eleven degrees Celsius, some twenty degrees below normal, with little sun. A camera inside the box saved an image every fifteen minutes. On the twentieth, Sheri Lubin, my assistant, who ran the nest box program and loved it, was reviewing the archive when she noticed the nestlings had stopped moving. She went back through the frames and reconstructed the parents' day: pinned to the nest all morning to keep the chicks warm, then away for hours in the afternoon hunting insects that were not flying. The chicks died of cold, or hunger, or both.

The sensors saw the cold. The camera saw the behavior. But it took a person to notice that something was wrong and ask why.

figure-1-1789145612.jpg
Nest box 8 through a full breeding cycle, as seen by the wired camera that watched it for a decade: nest building, laying, incubation, hatching, and the first days of nestlings. These sequences trained the classifiers in "Heartbeat of a Nest" (Ko et al., 2010) to detect a bird's presence and count eggs; even human experts found the eggs hard to count when feathers or nest material covered them. In the last hatching frame, the adult's tail fills the image as she leans out of the entrance hole. Reproduced from Ko et al. 2010.

Seven years later, the same nest box taught us what a machine could be trained to notice. Working with computer vision researchers at UCLA, we fed ten seasons of nest box images into classifiers that learned to tell whether a bird was present, count the eggs beneath it, and mark the transitions from laying to incubation to hatching. Five thousand images became a single trace, rising and falling with the parent's comings and goings, that looked so much like a cardiogram we titled the paper "Heartbeat of a Nest." The machine was right about four times in five.

And then the traces showed two arrhythmias. At box 8, a female kept incubating for weeks past the normal period. At another box, the whole cycle seemed to reset in mid-season. The algorithm flagged both. A person went back to the images and found the explanations: only one of four eggs had hatched and the mother would not give up on the rest; a mountain chickadee had built a nest, abandoned it, and been succeeded by a western bluebird. The machine found the irregular heartbeat. The biologist made the diagnosis.

That second step is the part of science the Nature feature says today's models lack: attention to the small discrepancy, the willingness to hold an anomaly in mind, the judgment that this one matters. Ball's sources call it abductive reasoning, the leap that invents a cause. Field ecologists call it Tuesday.

I do not think this gap is permanent. The researchers Ball quotes are careful to say the same. But it is where the line fell in 2010, it is where it falls today, and it tells us how to divide the labor.

Dreaming in NEON, revised

NEON is fragile. Its renewal comes up in 2027 or 2028, under an administration that has cut science budgets across the board, and the value of a thirty-year record is not linear in its length: a twelve-year record that stops loses most of what it was built to detect. The released data are archived and citable and will survive. The continuity may not.

The crowd layer is robust but uncalibrated. A birder's checklist tells you where birders went as much as where birds were.

What is missing is the middle: the connection between the eighty-one places where we know the truth and the continent where we have only the crowd. And eighty-one is exactly the right number for one job, which is learning that connection.

Here is the proposal, in plain terms. Around each NEON site, draw a circle. Inside it, compare the crowd's records, eBird checklists and iNaturalist photographs, to NEON's standardized bird counts and phenology surveys taken at the same time. Compare the free gridded weather to the tower's actual thermometers. What comes out is a set of correction factors: how much the crowd underreports, how far the grid misses the ground, species by species and region by region. Then apply those corrections everywhere else in the region, where there is no tower and never will be. This is how satellite imagery has been calibrated against field plots for fifty years. It is what WeatherNext 3 did with weather stations. Nobody has done it for the living layer, because the crowd platforms do not own the ground truth and the ground-truth facility was never asked to reach outward.

The prototype is modest: three NEON sites in the Pacific Northwest, Abby Road, Wind River, and McRae Creek at the H.J. Andrews Forest, compared against our own Willamette Valley records and the single climate logger on a fence post at Owl Farm. If it works, it extends naturally through the network of biological field stations, each of which has a location, a resident naturalist, some weather record, and its own cloud of local observers. Each station becomes a calibrated-by-proxy node. That is roughly what the campfire wanted, assembled from parts that already exist.

If NEON persists, the calibration improves every year. If it does not, the correction has already been learned and the crowd layer carries on. Either way the hedge is built from the thing the machines are actually good at.

The naturalist's share

I am seventy-two. I am out of the business of writing proposals for towers. For a hundred dollars a month I have access to a laboratory of language models that would have seemed like science fiction at that campfire, and I spend my mornings with it the way I once spent them with graduate students, arguing about what the data might mean.

The machine's share of the work is memory, arithmetic, and tirelessness. It will find the correlation in ten million checklists without complaint. My share, for now, is the thing that happened in nest box 8: deciding what to measure, noticing which outlier is not noise, and holding a question open across years while the pattern refuses to fit. That division is contingent. I expect it to shift within my lifetime, and I would like to be watching when it does.

The dream was never really about the sensors. It was about a continent that could notice itself changing, and people who would know what to ask when it did. The tools have run ahead of us again. The science is still ours to catch up.