Satellite pixels, search trends and storm forecasts live at different resolutions, on different clocks, in different silos. Earth AI turns each silo into a foundation model, shows that fusing their embeddings lifts almost every prediction it tries, and puts a Gemini agent on top that plans, calls the models, and answers crisis questions an analyst would spend a day joining by hand.
It is the evening of September 23rd, 2024. A hurricane named Helene is spinning up in the Gulf of Mexico, and an emergency manager in Florida needs one list before she goes home: every county with more than 20,000 people that sits inside the forecast band of hurricane-force winds. Not the whole state. Not a news article. A list with populations next to it, sorted, so that cooling centers, buses and water can be staged tomorrow morning.
Every piece of that answer exists somewhere. The cyclone forecast lives in a weather lab as fifty possible storm tracks. County boundaries live in a mapping service as polygons. Populations live in the Census Bureau, mirrored through a statistics service. None of those three systems knows the other two exist. So the analyst downloads the forecast, exports the polygons, pulls the census table, converts three coordinate systems into one, computes which polygons overlap which wind band, joins the population column, filters, sorts. If she is fast, that is an afternoon. Helene made landfall three days later.
That afternoon is the problem this paper is about. Geospatial data is not scarce. It is siloed: it comes in three families that were never designed to be read together.
Each lane below is one family of Earth data drawn at its native grain: the size of one cell is the size of one measurement. Pick a question and watch which lanes it needs and how many hand-made joins connect them. The counter is the analyst's afternoon.
The paper sorts questions like these into a hierarchy of three rungs, and the widget follows it. A descriptive question ("the hottest day in New York in July 2024") touches one silo and needs no join: look it up. An analytical question ("how many hospitals were in the storm band when Katrina came ashore") touches two silos and needs a spatial join: a polygon from one system intersected with points from another. A predictive question ("which Indian cities will have the most vulnerable people exposed to flooding by November 2027") touches all three and needs a join plus a model: nobody has that number in a table, it has to be forecast.
It is tempting to think the fix is one giant table. Put every pixel, every postal code and every forecast in one database and query it. The widget shows why that fails before it starts: the three families do not share a grain.
Imagery is measured per pixel, and a pixel can be anything from ten centimeters to ten meters wide depending on the satellite. A single city block is a few hundred pixels at one resolution and a few hundred thousand at another. Population is measured per administrative region: a postal code, a census tract, a county. Those are irregular polygons that were drawn for mail delivery and elections, not for science, and they change size by a factor of a thousand between Manhattan and Montana. Environment is measured on a regular grid with a clock attached: a forecast is a value per grid cell per hour, and it is reissued every few hours, so the same location has many competing values dated by when they were computed.
There is a second, quieter reason the silos never merged: sparsity. Labels are precious. The census reports household income per tract once every few years; a river gauge exists on some rivers and not others; a human-annotated satellite image of a flooded road is a rare thing. Any method that needs a large labeled dataset per question is dead on arrival. Whatever lives above the silos must be able to learn a new target variable from a few hundred labeled regions, not a few million.
Geospatial AI did not start with agents. It started with specialized models: one network for land cover, one for buildings, one for crop type, each trained on its own dataset and speaking to nobody. The next step was general-purpose foundation models for Earth observation, such as Prithvi and SatMAE, pretrained on large piles of multispectral satellite imagery so that one backbone could be fine-tuned for many tasks. That solved half of the imagery silo. It did nothing for population or environment, and it produced a backbone, not an analyst.
Most recently, the field started pointing large language models at geospatial data and building benchmarks to test them: MapEval for map reasoning, GeoChain for chain-of-thought geography. Those experiments exposed a gap that this paper is built around. A language model reasons well over natural-language facts it has been handed. It reasons badly over a polygon, a raster, or a table of forecasts it has to go find, parse and join itself.
Hold onto one number for the rest of the lesson. When the authors predict FEMA's twenty disaster-risk indices for every census tract in the country, using landscape embeddings alone gets an average R² of 0.54. Using population embeddings alone also gets 0.54. Using both gets 0.60. The two silos know different things about the same tract, and the whole paper is an argument that this is true everywhere: imagery knows the physical hazard, population knows who can absorb it, environment knows when it arrives.
The paper's own one-sentence summary is buried at the end of its architecture section, and it is the sentence to remember: capable models distill reality, and a powerful agent reasons over that distillation. Two layers, with a deliberate division of labor between them.
The bottom layer is three families of foundation models, one per silo. Each takes its silo's raw, incompatible unit (pixels, polygons, grid-cell-hours) and produces something a downstream system can actually use: a fixed-length embedding vector, a detection, a classification, or a forecast with uncertainty attached. That is the distillation. Reality goes in at whatever grain it has; a compact, comparable representation comes out.
The top layer is a single Gemini-powered agent that never touches a pixel. It reads the question, decomposes it into steps, decides which model or tool each step needs, calls them, checks the result, and writes the answer. That is the reasoning. The agent's competence does not come from knowing geography; it comes from having a menu of specialists and knowing how to sequence them.
This is the paper's Figure 1 rebuilt so you can interrogate it. Click any box to see what goes in, what comes out, and what it is trained on. Watch the data types change across the diagram: raw pixels become vectors, vectors become numbers, numbers become a sentence with a map attached.
Read the diagram left to right and notice that the arrows carry different kinds of things. The leftmost arrows carry raw data: a stack of satellite tiles, a table of Search Trends per postal code, a gridded weather forecast. The middle arrows carry distillations: a 64-number embedding per ten-meter pixel from AlphaEarth, a fixed-length vector per postal code from Population Dynamics Foundations, fifty candidate storm tracks from the cyclone model. The rightmost arrows carry arguments: a plan, a tool call, a table, a sentence. The layers are separated by the type of thing that flows between them, not by an org chart.
Why not train one enormous model on everything and ask it questions directly? The paper gives three reasons, and each one is visible in the diagram.
First, label scarcity lives at the bottom. A foundation model per silo can be pretrained once, self-supervised, on the ocean of unlabeled data each silo has (300 million satellite images; every postal code's search behavior), and then adapted to a new target variable with a few hundred labels. If the agent had to learn geography from question–answer pairs, it would need labels that do not exist.
Second, complementarity lives in the middle. Because each silo's model produces an embedding in its own space, an analyst (or the agent) can concatenate two of them and train a tiny model on top. Chapter 7 is entirely about this: the same boosted-tree recipe, fed landscape embeddings plus population embeddings, beats either alone on 39 of 41 target variables.
Third, extensibility lives at the top. The agent's capabilities are organized by domain, as tools or "expert" sub-agents. Adding a new model means adding a new tool; the reasoning layer does not retrain. The authors say this explicitly: the domain-oriented design exists so they can plug in new geospatial domains and swap implementations as language models and frameworks improve.
| Finding | Evidence in the paper | Chapters |
|---|---|---|
| State-of-the-art models per silo | Remote Sensing Foundations: best zero-shot classification and retrieval on most public benchmarks; 31.83% mAP open-vocabulary detection on DOTA versus 13.77% for the base model. Population Dynamics Foundations: R² 0.85 across 17 countries; independently validated by seven outside groups. | 2, 3, 4, 5, 6 |
| Synergy across silos | Fusing AlphaEarth and Population Dynamics embeddings lifts FEMA-risk R² by 11% on average and beats either alone on all 21 CDC health variables; weather covariates cut cholera forecast error by a third; a cyclone ensemble plus population embeddings predicted Hurricane Ian's damaged-building count within 3% three days out. | 7, 8 |
| Agentic reasoning | The Geospatial Reasoning Agent scores 0.82 on a 100-question fill-in-the-blank benchmark versus 0.50 for Gemini 2.5 Pro with search, maps and code tools; 0.87 versus 0.38 on ten rubric-graded crisis scenarios. | 9, 10 |
The paper reads like the third step of a sequence you can reconstruct. Google had already shipped the pieces separately: a building-detection model, a flood-forecasting system, a population-embedding model released in 2024, a cyclone model in 2025, AlphaEarth in 2025. Each team benchmarked its own model on its own tasks. The first question someone asked was probably the cheap one: "if I have the population embedding and the AlphaEarth embedding for the same tract, does stacking them help?" It did, by 11%. The second question follows immediately: "if the models are complementary, who is going to do the stacking for a user who is not a data scientist?" The answer to that is an agent, and the benchmark of ten crisis scenarios is what you build when you want to prove the agent can do the stacking on its own.
Start with the imagery silo, because it is the one where the paper's models are most concrete and its benchmarks most public. The first capability the authors want is simple to state and was, until recently, impossible: type "a solar farm next to farmland" and have a model find every satellite tile that matches, with no training examples of solar farms at all.
The way to get there is a vision–language model in the CLIP family: two encoders, one for images and one for text, trained so that an image and its caption land close together in a shared embedding space and an image and a stranger's caption land far apart. Once that space exists, classification is free: embed the tile, embed a handful of candidate labels as sentences, and pick the label whose vector is closest. Retrieval is the same trick run the other way.
Six synthetic overhead tiles on the left, six text prompts across the top, and the cosine similarity between every pair in the grid. The base encoder was trained on ordinary web photos, so it half-knows what a runway looks like from above. Toggle the remote-sensing fine-tune and watch the diagonal sharpen: that is what 18 million caption pairs of overhead imagery buy. Click a prompt to run retrieval; click a tile to run zero-shot classification.
The widget is a cartoon of the real thing, but the cartoon is faithful in one respect: a general-purpose model is not wrong about overhead imagery, it is blurry. Off-the-shelf SigLIP2 gets 72.40% top-1 zero-shot accuracy on RESISC45, a 45-class scene dataset. That is far above chance. The remote-sensing version of the same architecture, trained on the same recipe plus the datasets below, gets 80.13%. On the harder FMoW benchmark the jump is 41.25% to 48.13%. Nothing about the architecture changed. Only the captions did.
The bottleneck for a remote-sensing VLM is not images. Google has decades of aerial and satellite imagery. The bottleneck is text: nobody writes captions for satellite tiles. The paper's answer is three datasets, each built by a different trick.
Two model families are fine-tuned on those datasets: MaMMUT (a joint encoder–decoder design) on RS-Landmarks plus the smaller RS-WebLI, and SigLIP2 (a sigmoid-loss contrastive model) on the larger RS-WebLI plus RS-Landmarks plus RS-Global. Both use 400-million-parameter image encoders and 400-million-parameter text encoders, sized so that the models can run on constrained hardware, not just in a datacenter.
Zero-shot classification asks: given a tile and the names of the 45 (or 30, or 21) classes as text, does the nearest class embedding match the ground truth? Retrieval asks the symmetric questions: given a caption, is the right image among the top results (text-to-image), and given an image, is the right caption (image-to-text)? The paper reports the average of top-1, top-5 and top-10 for retrieval.
| Zero-shot top-1 accuracy | FMoW | SkyScript | RESISC45 | UCM | AID |
|---|---|---|---|---|---|
| SigLIP2 400M (base) | 41.25 | 65.94 | 72.40 | 81.43 | 75.56 |
| RS-SigLIP2 400M | 48.13 | 68.13 | 80.13 | 84.86 | 78.26 |
| RS-MaMMUT 400M | 47.24 | 69.46 | 72.31 | 80.29 | 71.96 |
| RemoteCLIP (prior RS model) | – | – | 79.84 | – | 91.30 |
| Retrieval, avg of top-1/5/10 | RSICD I2T / T2I | UCM-Captions I2T / T2I | RSITMD I2T / T2I | NWPU I2T / T2I |
|---|---|---|---|---|
| SigLIP2 400M (base) | 27.36 / 27.93 | 70.48 / 68.52 | 24.62 / 25.14 | 31.86 / 36.70 |
| RS-SigLIP2 400M | 38.37 / 37.64 | 76.67 / 75.33 | 43.14 / 47.26 | 45.12 / 37.74 |
| RS-MaMMUT 400M | 33.33 / 33.59 | 74.76 / 71.79 | 42.63 / 42.58 | 41.44 / 32.28 |
Read the tables honestly, the way a reviewer would. On four of five classification benchmarks the remote-sensing fine-tune beats its own base model, and on three of them it beats every prior remote-sensing model. But RemoteCLIP still wins AID by a wide margin (91.30 versus 78.26), and the SkyScript number carries an asterisk because that model was evaluated in-domain. Retrieval is the cleaner story: RS-SigLIP2 wins seven of eight columns, and the gain on RSITMD (24.62 to 43.14 image-to-text) is close to a doubling. The paper's summary phrase is "state of the art on the vast majority of benchmarks," and the tables support exactly that phrase, no stronger.
A tile x of shape [3, 224, 224] goes through the image encoder and comes out as a single vector v of a few hundred to a thousand numbers, then normalized to unit length. A prompt string goes through the text encoder and comes out as a unit vector u. The score is their dot product, which for unit vectors is the cosine of the angle between them:
Zero-shot classification over K class names is K dot products and an argmax. Retrieval over a million tiles is a million dot products, which is why the whole design cares about pre-computing and caching image embeddings: encode every tile once, then any new question is a matrix multiply against the cache. That same caching idea reappears in the detection model of the next chapter and in the agent's imagery toolset in Chapter 9.
# zero-shot classification with a contrastive VLM (shapes in comments) v = normalize(image_encoder(tile)) # [D] one unit vector per tile U = normalize(text_encoder(["a photo of a %s" % c for c in classes])) # [K, D] scores = U @ v # [K] K cosines label = classes[scores.argmax()] # retrieval: same math, cached the other way round V = normalize(image_encoder(all_tiles)) # [N, D] computed ONCE, stored u = normalize(text_encoder("a solar farm next to farmland")) top = (V @ u).topk(10) # [10] a million dots, no re-encoding
Classifying a whole tile is a coarse tool. The crisis-response prompts in this paper ask things like "which schools and grocery stores in Lahaina were in the burned area" and "how many lakes are in the residential parts of this zip code." Those need boxes, not labels: where, in the image, is each object of a kind I name in words. That is open-vocabulary object detection (OVD), and cataloging every possible class up front is hopeless for the planet, so the vocabulary has to be open.
The base detector is OWL-ViT v2, the same idea as the VLM but with a twist: the image encoder produces one embedding per candidate box instead of one per image, and the text encoder produces one embedding per query. A box is a detection of "grain silo" if its embedding is close to the embedding of the words "grain silo." Because the two encoders are independent, the box embeddings for a region can be computed once and cached, and any new query is a dot product against the cache. The paper calls this split processing and it is the reason the agent can re-query the same satellite tiles cheaply.
Zero-shot OWL-ViT v2 gets 13.77% mean average precision on DOTA and 14.98% on DIOR, two standard aerial detection benchmarks. Overhead objects are small, oddly oriented and packed together; a detector trained on web photos of dogs and bicycles half-recognizes them. The remote-sensing version, RS-OWL-ViT-v2, more than doubles both numbers to 31.83% and 29.39%, with a two-stage cooldown fine-tune:
Why that order? Because the noisy pseudo-boxes teach what overhead objects look like in general (a thing with a shadow, a rectangle at a strange angle), while the human labels teach precision. Running them the other way round would let 3M noisy images wash out the precise ones.
Even a good zero-shot detector fails on ambiguity. "Silo" could mean a grain silo or a missile silo; "boat" at ten-meter resolution is a few pixels. The classic fix is few-shot learning: give the detector a handful of labeled examples of the thing you mean. The expensive part is choosing those examples: a human has to annotate them, so you want the ones that teach the most per click.
FLAME (from the same group) is a one-step active-learning strategy for exactly that. It runs the zero-shot detector, looks at the embeddings of all the candidate boxes it produced for your query, and picks the examples to show a human in two moves.
Each dot is one candidate box the detector found for a query, drawn in a 2-D projection of its embedding. Dense clumps are boxes the model is confident about (many look alike). Step 1 keeps the low-density margins: the ambiguous ones. Step 2 clusters those with k-means and takes the point nearest each center, so the human labels examples that are both uncertain and different from each other. Slide the budget to see how few labels FLAME needs; the mAP readout follows the paper's numbers.
Step 1, uncertainty filtering. Combine each box's image features with the text query into one embedding, reduce its dimensionality, and estimate the density of the resulting cloud. Points in dense regions are boxes the model has seen many near-copies of; it is sure about them. Points on the low-density margins of the core are the ambiguous ones. Keep the margins.
Step 2, diversity sampling. The margins may still contain fifty near-duplicate ambiguous boxes. Run k-means on them with k equal to your labeling budget and take, from each cluster, the sample closest to its center. Now the human sees k examples that are uncertain and spread out.
With 30 labeled examples per category, this lifts mAP on DOTA from 31.83% (zero-shot) to 53.96%, and on DIOR from 29.39% to 53.21%. That beats SIoU, a recent few-shot detection method built specifically for small objects, on both benchmarks (45.88% and 52.85%). Thirty clicks per class, chosen well, nearly doubles the detector.
# FLAME in ten lines (the paper's two moves) E = embed_boxes(image_feats, text_query) # [N, D] one vector per candidate box Z = reduce(E, dims=2) # [N, 2] cheap to estimate density in dens = kde(Z) # [N] how crowded is each point's neighbourhood margin = Z[dens < quantile(dens, 0.3)] # the uncertain 30% centers, assign = kmeans(margin, k=budget) picks = [margin[assign == j][argmin(dist(margin[assign == j], centers[j]))] for j in range(budget)] # show `picks` to a human; train the lightweight few-shot classifier on their labels
The VLM and the detector both answer questions in words. Plenty of geospatial work does not: segmenting a flooded road pixel by pixel, outlining every building in a city, classifying scenes for a dataset that will be fine-tuned anyway. For those, what you want is a backbone: a vision transformer whose features are so good that a small task head on top, trained on a few thousand labels, beats a model trained from scratch on millions.
The paper's backbone recipe, called RS-Global MTP, has two stages, and the second one is the unusual part.
Take the image-only version of RS-Global and scale it ten times, to 300 million aerial and satellite images, randomly cropped and resized during training so that the effective resolution varies continuously from 10 cm to 10 m per pixel. Then train a masked autoencoder (MAE): chop each image into patches, hide most of them, and ask the network to reconstruct the hidden pixels from the visible ones. No labels anywhere.
Left: an 8×8 patch grid of a synthetic overhead scene. Drag the mask ratio and press Reconstruct to watch the encoder fill in what it never saw: the only way it can is by having learned what fields, roads and roofs look like. Right: Stage 2's schedule. Multi-task pretraining does not sum three losses in one step; it alternates, one task per step, so the encoder hears one teacher at a time.
Why does hiding three quarters of an image teach anything? Because the reconstruction target is not a class name, it is the world itself. To fill in a hidden patch of farmland, the encoder must have learned that furrows continue in straight lines, that a road that enters a tile usually exits it, that a roof has a shadow on the side away from the sun. Those are the same regularities every downstream task needs. The paper credits MAE with capturing "low-level spectral patterns as well as higher-level spatial structures unique to remote sensing data."
MAE features are general but not task-shaped. So the second stage adds three supervised objectives on internal labeled data: scene classification, semantic segmentation, and object detection. The naive way to combine them is to compute all three losses on each batch, add them up, and take one gradient step on the sum. The paper does something else: Alternating Gradient Descent (AGD), where each training step optimizes one task, cycling through them.
The reason is task interference. Segmentation wants the encoder to keep fine boundaries; classification wants it to throw boundaries away and summarize; detection wants something in between. Summing the losses makes every step a compromise none of the three tasks asked for, and the gradients can partially cancel. Alternating lets each task's head and its share of the encoder specialize for one step while the shared representation still has to serve all three over time. The paper's phrase: it "mitigates task interference, allowing different task experts to specialize while still contributing to a unified representation."
The backbone was evaluated by fine-tuning it on 13 downstream benchmarks across four task types, always with the encoder weights tunable, the same batch size and epochs for every model, and only the learning rate tuned per model. Results are reported relative to a ViT-L/16 pretrained on ImageNet, the standard off-the-shelf starting point:
| Backbone (relative to ImageNet ViT-L/16) | Semantic seg. (mIoU) | Detection (AP50) | Classification (acc.) | Instance seg. (mAP) |
|---|---|---|---|---|
| Random init | −16.23% | −17.01% | −24.14% | −40.87% |
| DINOv2 ViT-L/14 | +4.44% | +4.69% | +4.89% | +75.94% |
| SigLIP2 So400m | +4.50% | +2.47% | +4.74% | +63.75% |
| RS-SigLIP2 (the Chapter 2 model) | +4.15% | +2.52% | +4.79% | +66.74% |
| RS-Global MTP | +6.09% | +5.91% | +4.58% | +82.92% |
Three things to notice. The instance-segmentation column is enormous for every pretrained backbone, which tells you the ImageNet baseline is simply bad at outlining buildings; the interesting comparison is RS-Global MTP against DINOv2, and the remote-sensing backbone wins by seven points there. On classification the remote-sensing backbone is slightly behind DINOv2 (4.58% versus 4.89%), which the paper does not hide. And the VLM from Chapter 2, when used as a backbone, is no better than base SigLIP2: contrastive image–text training makes a great retriever and an ordinary backbone. Different objectives, different strengths. Averaged over all task categories, RS-Global MTP improves 14.93% over the ImageNet baseline.
Against published state of the art on public splits the backbone also leads: 81.70% on FMoW classification (prior best 79.30%), 65.72% mIoU on FLAIR segmentation (63.10%), 85.50% AP50 on DIOR detection (81.10%).
# Stage 2 in code: alternate, do not sum tasks = [("cls", cls_loader, cls_head), ("seg", seg_loader, seg_head), ("det", det_loader, det_head)] for step in range(total_steps): name, loader, head = tasks[step % 3] # one task per step, round robin x, y = next(loader) feats = encoder(x) # shared ViT, MAE-initialised loss = head.loss(feats, y) # ONLY this task's loss loss.backward(); opt.step(); opt.zero_grad()
Now the second silo, and the one with no pixels in it at all. What does a postal code look like to a model? Not from above: from the inside. What businesses are there, what people search for, how busy the streets are at noon, what the weather was. The Population Dynamics Foundations model (PDFM) turns all of that into one vector per region, and that vector turns out to predict things nobody put into it: diabetes rates, home values, night-time lights, dengue.
Four families of signal, all aggregated to an administrative region (a US postal code, or a similarly sized region elsewhere), all privacy-preserving by construction because nothing below the region level ever exists in the model:
| Signal | Source | What it encodes |
|---|---|---|
| Built environment | Google Maps | Counts and kinds of places: restaurants, clinics, schools, parks, road density |
| Human behavior | Search Trends | Which topics people in the region search for, relative to elsewhere |
| Human presence | Anonymized busyness | How occupied places are through the day and week |
| Natural environment | Weather and air quality | Temperature, precipitation, pollution over the period |
The Search Trends channel is the clever one, and the global version of the model changed how it is built. Instead of counting the top raw queries per region (which are language-specific and privacy-sensitive), the global model counts the top Knowledge Graph entities that the queries resolve to. "Taylor Swift boyfriend" and "KC tight end" both map to the entity Travis Kelce. For each region, the model records the frequencies of the top 500 entities searched at least 20 times; entities are ranked by how many regions they appear in across 17 countries, and the top 1,000 form a fixed vocabulary. The least popular entity in that vocabulary still appeared in searches from more than 9,000 regions, which is the privacy floor: nothing rare enough to identify anyone survives.
Regions are not independent. A postal code shares a border, a river, a commute pattern with its neighbors. So the model is a graph neural network: each region is a node holding its four signal vectors, edges connect nearby regions, and a few rounds of message passing let each node's embedding absorb its neighborhood. The training objective is self-supervised (the original PDFM paper trains the GNN to reconstruct held-out signals from neighbors), so the embedding is a compressed summary of "what this region and its surroundings are like" that was never told about any downstream target.
A hex map of synthetic regions, each with four signal channels (Maps, Search, busyness, weather) shown as its color mix. Press Message pass to watch each node blend in its neighbors: that is the GNN, and the result is the embedding. Then Hide 20% of the regions' target value and compare two ways of guessing it: inverse-distance weighting (the classic GIS method, which only knows where a region is) versus nearest neighbors in embedding space (which knows what a region is like). The R² readouts are computed live on this toy; the paper's real numbers are in the table below.
The toy makes the paper's central claim about this silo visible. Inverse-distance weighting assumes that nearby regions are similar. Embedding-space neighbors assume that similar regions are similar, wherever they are: a university town in Ohio has more in common with a university town in Oregon than with the farmland next door. When the hidden target correlates with what a region is like rather than where it sits, the embedding wins.
The 2024 PDFM covered the contiguous United States for a single year. This paper adds two things.
Global coverage, 17 countries. Australia, Belgium, Brazil, Canada, France, Germany, India, Italy, Japan, Mexico, the Netherlands, Nigeria, Portugal, Spain, Switzerland, the United Kingdom and the United States. The embeddings are built in one shared space, so a model trained on US labels can be applied in the UK or Brazil. The benchmark is spatial interpolation: hide 20% of a country's regions, predict four globally available targets from Earth Engine (elevation, night-time lights, population density, tree cover), score with R².
| Global PDFM, R² (hold out 20% of regions) | Elevation | Night lights | Pop. density | Tree cover | Mean |
|---|---|---|---|---|---|
| United States | 0.98 | 0.93 | 0.88 | 0.87 | 0.92 |
| India | 0.97 | 0.90 | 0.87 | 0.85 | 0.90 |
| Japan | 0.91 | 0.94 | 0.94 | 0.81 | 0.90 |
| Netherlands | 0.87 | 0.88 | 0.89 | 0.42 | 0.77 |
| Nigeria | 0.79 | 0.52 | 0.52 | 0.71 | 0.63 |
| All 17 countries | 0.91 | 0.88 | 0.88 | 0.73 | 0.85 |
Notice where it is weak, because the weaknesses are informative. Tree cover in the Netherlands (0.42) and Switzerland (0.50) is hard because those countries have little of it and what exists is not where the searches and businesses are. Nigeria is the lowest overall (0.63) with night lights and population density at 0.52: the input signals are sparser where Maps coverage and search volume are lower. The embedding is only as rich as the digital footprint it summarizes.
The stronger test is cross-country extrapolation: train on regions in Belgium, Switzerland, Germany, Spain, the UK, Italy and the Netherlands, predict France, using Eurostat's NUTS-3 statistics. Against the two standard GIS interpolators:
| Predicting France from 7 neighbors | GDP/capita R² | MAPE | Death rate R² | MAPE | Fertility R² | MAPE |
|---|---|---|---|---|---|---|
| Inverse Distance Weighting | 0.00 | 28.03% | −0.07 | 19.29% | 0.51 | 8.78% |
| Radial Basis Functions | 0.13 | 25.71% | 0.52 | 12.91% | 0.50 | 9.35% |
| PDFM embeddings | 0.52 | 19.03% | 0.78 | 7.99% | 0.71 | 6.79% |
Read the first row. Inverse-distance weighting gets an R² of exactly zero on GDP per capita, and a negative R² on death rate: it is worse than predicting the mean. That is what happens when you extrapolate into a country using only the coordinates of regions in other countries. The embedding does not know France is France; it knows a NUTS-3 region full of pharmacies, ski searches and low busyness in August looks like certain Swiss and Italian regions, and it borrows their numbers.
Temporal embeddings, monthly since July 2023. The static embedding is a snapshot. The temporal version rebuilds it every month from that month's signals only. Tested on monthly emergency-department visits for COVID-19, flu and RSV in a fixed holdout of ten US states over two years, the time-matched embedding consistently reduces mean absolute error compared with a static embedding frozen in July 2023, with the largest gains in fall and winter when disease burden is highest. That is the "Temporal drift" toggle in the widget: a static snapshot slowly goes stale as behavior shifts, and a monthly embedding tracks it.
Foundation-model papers usually benchmark themselves. This one also reports independent validations of the static US embeddings by groups that had no stake in the result, and the table is worth its space because the tasks are so different from each other:
| Who | Task | Result |
|---|---|---|
| CARTO | Interpolate their proprietary Human Activity Index across the US, random forest on PDFM | R² 0.882 on the test split |
| CARTO | Iowa liquor sales | PDFM alone R² 0.666 (0.097 above demographics alone); ensemble with demographics 0.707 |
| SustGlobal | Home-insurance premium tier, XGBoost on PDFM | 93% top-2 accuracy at county level, 89.5% at zip level |
| University of Oxford | 1–12 month dengue forecasting in Brazil | Best model R² 0.849 → 0.901 with PDFM; with TimesFM at 6 months 0.500 → 0.686, at 12 months 0.456 → 0.656 |
| Institute for Disease Modeling | Poliovirus risk in Nigeria, ridge regression with spatial effects | Better fit (DIC/WAIC); predicted risk changed more than 2-fold in 18% of local government areas |
| Mount Sinai / Boston Children's / Harvard | Measles vulnerability maps | Super-resolved from county level to zip-code level |
The third silo is the one with a clock. Imagery and population describe the world as it is or was. Environment models describe what it will be, with uncertainty attached, and they are reissued constantly. The paper does not train new environment models; it integrates three that Google already runs, and what matters for the rest of the lesson is the shape of what each one returns.
| Model | What it returns | Horizon and history |
|---|---|---|
| Weather (Maps Platform Weather API, MetNet-family nowcasting) | Hourly temperature, precipitation, wind, UV per location | Hourly to 240 hours, daily to 10 days |
| Floods (Flood Forecasting API) | Riverine flood predictions from gauge data: inundation area, severity level, probability | Current, plus history back to August 1st, 2025 |
| Cyclones (experimental, stochastic neural networks) | Formation, track, intensity, size and shape as 50 possible scenarios | Up to 15 days ahead; history back to January 1st, 2022 |
The cyclone model deserves the widget, because "50 possible scenarios" is the most important data type in this paper's crisis chapters and the least intuitive one.
A storm starts at the dot on the left. Each thin line is one ensemble member: a plausible track sampled from the model. Slide the lead time and watch the fan open. The shaded band is the union of every member's hurricane-force wind radius, and the grid behind it is a state of counties with populations. The readout counts counties over 20,000 people that any member's wind band touches: the exact computation the agent performs for Hurricane Helene in Chapter 9.
Two things the widget teaches. First, an ensemble is not a single answer with error bars; it is fifty complete, internally consistent futures. Any question you ask of the storm ("which counties get hurricane-force winds") can be asked of each member separately, which is how the damage model in Chapter 8 produces a distribution of damaged buildings rather than one number. Second, the fan widens with lead time, so the same question has a very different answer at three days and at fifteen. The agent has to carry the "as of" timestamp through every step, and the crisis prompts in Chapter 10 all pin one.
# what each environment tool hands back to the agent (schematic) weather = {"location": (lat, lon), "hourly": [{t, temp_c, precip_mm, wind_kph}] * 240} flood = {"gauge": id, "severity": "high", "probability": 0.71, "inundation": Polygon} cyclone = {"init": "2024-09-23T12:00Z", "members": [ # 50 of these {"track": [(lat, lon, t)], "wind_kt": [...], "r_hurricane_km": [...]}]} # the question "counties in hurricane-force winds" against an ensemble band = union(buffer(m.track, m.r_hurricane_km) for m in cyclone.members) # one polygon hit = [c for c in counties if c.geometry.intersects(band) and c.pop > 20000]
This is the chapter the whole paper turns on. Each silo's model is good. The claim is that they are complementary: that a landscape embedding and a population embedding for the same census tract know different things, so a model fed both beats a model fed either. The evidence is two tables covering 41 target variables, and the recipe that produced them is simple enough to rebuild in an afternoon.
Take two public embeddings. AlphaEarth Foundations summarizes optical satellite imagery, radar, climate simulations and more into a 64-band embedding at 10-meter resolution, one image per year, available in Earth Engine. Population Dynamics Foundations gives one vector per US postal code. Neither is at census-tract resolution, and tracts are what FEMA and the CDC report on. So both have to be moved:
One tract, one AlphaEarth band. The pixels inside the tract boundary each carry a value for this band; press Histogram to bin them into ten fixed buckets. That row of ten counts is what the boosted trees see for this band. Step through bands to fill the 64×10 feature block on the right.
Why a histogram and not a mean? Because FEMA risk is about extremes. Coastal-flooding risk depends on how much of a tract is at sea level, not on the tract's average elevation. A mean of the elevation-like bands would say "moderately low"; a histogram says "12% of pixels in the lowest bin," which is the number that matters. The feature construction is a modeling decision disguised as plumbing.
Every number here is from the paper's Tables 9 and 10: R² on census tracts in the 20% of held-out counties. Pick a target, toggle feature sets, and watch which silo carries which variable. Then press Sweep all to see the whole picture at once: where fusion wins, by how much, and the two places it does not.
On the twenty FEMA National Risk Index targets, PDFM alone averages R² 0.54, AlphaEarth alone averages 0.54, and both together average 0.60, an 11% relative lift. But the average hides the structure, and the structure is the lesson. Sort the targets by which silo wins alone:
| FEMA target | PDFM | AlphaEarth | Both | Who knows this? |
|---|---|---|---|---|
| Earthquake risk | 0.72 | 0.84 | 0.86 | The landscape: faults, terrain, building density from above |
| Social vulnerability score | 0.30 | 0.44 | 0.48 | Hard for both; imagery surprisingly ahead |
| Drought risk | 0.67 | 0.50 | 0.67 | Population: agriculture, search patterns; imagery adds nothing |
| Wildfire risk | 0.77 | 0.54 | 0.76 | Population wins; fusion slightly hurts |
| Coastal flooding risk | 0.33 | 0.24 | 0.38 | Both weak, fusion best: coast + who lives there |
| Overall risk score | 0.53 | 0.49 | 0.60 | The composite gains most from fusion |
FEMA's own definition explains the pattern. A risk index multiplies expected annual loss (a physical quantity: how likely and how severe the hazard is, which imagery, terrain and coastline encode) by social vulnerability and divides by community resilience (human quantities: income, age, density, which the population embedding encodes). Risk is bimodal by construction, so a model that sees only one mode is capped. The overall Risk Score, the most composite target, gains the most from fusion: 0.53 and 0.49 alone, 0.60 together.
On the 21 CDC health indicators the story is cleaner: fusion wins on every variable. The means are 0.42 for PDFM alone, 0.56 for AlphaEarth alone, 0.60 for both.
A gradient-boosted tree ensemble fits a sequence of small decision trees, each one correcting the residual error of the ones before. It is the standard tool for tabular data with hundreds of numeric features and a few thousand rows, which is exactly this shape (about 700 features, tens of thousands of tracts). It needs no feature scaling, handles the histogram bins and the embedding dimensions in the same way, and gives a direct measure of which features it split on.
The fixed hyperparameters are the point. The authors could have tuned per target and squeezed out a few more points of R². Instead they held everything constant across 41 targets so that the only variable is the feature set. The result "suggests that a simple modeling setup can benefit from the gains of combining embeddings without tuning to each individual label," which is the property an agent needs: in Chapter 9, the Spatiotemporal Model Training expert trains models like this on the fly for whatever variable a user names.
# the whole synergy experiment, in the shape the paper describes ae = alphaearth_histograms(tracts, year=2023, bins=10) # [T, 640] 64 bands x 10 bins pd = area_weighted_pdfm(tracts, overlap_2020) # [T, D] zip vectors -> tract X = {"pdfm": pd, "alphaearth": ae, "both": np.hstack([ae, pd])} train_ct, test_ct = split_counties(counties, test_frac=0.2) # split by COUNTY, not by tract for target in fema_20 + cdc_21: for name, feats in X.items(): m = GradientBoostingRegressor(**FIXED) # same settings for all 41 x 3 fits m.fit(feats[in_(train_ct)], y[target][in_(train_ct)]) r2[target][name] = m.score(feats[in_(test_ct)], y[target][in_(test_ct)])
Chapter 7 fused two static embeddings to predict a static number. The environment silo adds time, and time changes the question from "what is this place like" to "what happens here next week." Two case studies show the same move with a clock attached: take a forecast, and give it what the other silos know.
Bellwether, a Google X project, built a model that answers one question per building: will this building be wind-damaged by this storm? The training data is 1.5 million data points across tens of thousands of residential and commercial properties, six years (2017–2022), 22 hurricanes, with damage labels and property details (material, roof type), plus storm information from the HURDAT2 archive, building heights from Open Buildings, and the Population Dynamics embedding of the building's region. Six of the most recent storms are held out. The model is, again, gradient-boosted trees.
Here is the trap in that dataset, and the paper walks straight into it on purpose: fewer than 4% of properties are damaged in any storm. A model that says "undamaged" to everyone scores 96% accuracy. So the paper reports the numbers that matter, the ones for the minority class:
| Held-out test, observed tracks | Recall | Precision | F1 |
|---|---|---|---|
| Class 0: undamaged buildings | 0.98 | 0.98 | 0.98 |
| Class 1: damaged buildings | 0.60 | 0.57 | 0.59 |
Recall 0.60 means the model finds six of every ten buildings that will actually be damaged; precision 0.57 means when it flags a building, it is right a little more than half the time. Modest, honestly reported, and useful, because the operational question is not "which building" but "how many," and counts average out per-building errors.
Then comes the forecast. Replace the observed track from HURDAT2 with the cyclone model's forecast track, issued September 26th, 2022 at 00 UTC, three days before Ian's landfall. There are fifty forecast members, so run the damage model fifty times and collect fifty counts.
Each ensemble member is pushed through the per-building damage classifier, producing one predicted count of damaged buildings. Press Run ensemble to build the distribution member by member. The dashed line is the ground truth for this dataset, 2,496 buildings; the mean of the fifty forecasts landed at 2,575. Below, the fifty precision/recall/F1 values for the damaged class spread into box plots.
2,575 predicted against 2,496 observed is a 3% error, three days before landfall, and the spread of the distribution tells a relief planner how much to hedge. The paper is careful here: this is one storm at one lead time, the dataset is a subset of buildings in the affected zone (the true damage count was far higher), and some "wind damage" labels are probably flood damage. Read it as a demonstration of the shape of the pipeline, not as a validated forecasting product. The shape is the point: an ensemble forecast, a per-unit model fed static embeddings, and a distribution out.
The Democratic Republic of Congo has endemic cholera, and with WHO's African regional office the authors built a sub-national forecaster for weekly case counts in 24 provinces. The base model is TimesFM 2.0, a decoder-only foundation model for time series that can forecast a new series zero-shot, with no training on it. The experiment adds covariates in layers:
A synthetic weekly case series for one province with a rainy-season swell. Toggle the layers: TimesFM alone extrapolates the history; adding the weather forecast (rain and temperature drive transmission) lets it anticipate the swell; adding the province's PDFM embedding shifts the level to match what the province is like. The error readouts follow the paper's reported improvements.
| Layer | What it adds | Reported effect |
|---|---|---|
| Prophet / SARIMAX | Classical statistical baselines | The reference |
| TimesFM 2.0 zero-shot | Foundation-model extrapolation of the weekly counts alone | 35.4% lower RMSE than Prophet at a 20-week national horizon |
| + dynamic weather covariates | Forecast precipitation and temperature, the known drivers of transmission | A further 32.7% improvement |
| + PDFM static covariates | One embedding per province, describing what it is like | 5% lower RMSE at the regional level versus TimesFM alone |
The ordering of the gains is instructive. The biggest single step is swapping a classical model for a time-series foundation model, before any Earth AI signal is involved. The second biggest is the environment silo: weather is causal for cholera, so a forecast of rain is a forecast of cases. The population embedding adds the least, 5%, and only regionally, which makes sense: a static description of a province helps set its level but cannot tell you about next month's outbreak. Each silo contributes what it knows, and the paper reports the small gain as a small gain.
Everything so far is a specialist. Chapter 9 is the generalist that hires them. The Geospatial Reasoning Agent is built on Google's Agent Development Kit (ADK) with Gemini 2.5 Flash and Pro, and its design is a registry of capabilities organized by domain, plus a loop that plans, acts and checks. It is worth being precise about what is a tool and what is an expert, because the paper found the difference matters.
Some domains "succeeded as a simple list of tools." Others "needed more steering and guidance," so they were wrapped as expert sub-agents: a sub-agent has its own prompt, its own judgement about which of several data sources to use, and returns a synthesized answer rather than a raw payload.
| Capability | Kind | What it does |
|---|---|---|
| Demographics | Expert | Chooses and retrieves the right statistical variables from Data Commons for a natural-language question |
| Weather | Expert | Fetches and interprets historical (ERA5 via Earth Engine) and forecast weather, cyclones and floods |
| Spatiotemporal Model Training | Expert | Trains lightweight models on the fly on PDFM and TimesFM for interpolation, super-resolution and forecasting of a user-named variable |
| Earth Engine | Expert | Finds the right public dataset and band in the Earth Engine catalog by searching it, then fetches |
| Search | Expert | A fallback: Vertex AI grounding with Google Search when nothing specialized has the fact |
| Location & Places | Toolset | Turns names (a county, "hospitals in Louisiana") into polygon geometries via the Places and Places Aggregate APIs; the foundation for every spatial intersection |
| Remote Sensing & Imagery | Toolset | The Chapter 2–3 models: open-vocabulary classification, retrieval and detection, with cached image embeddings so repeat queries over the same area are cheap |
| General | Tools | Code generation for custom analysis, Earth Engine access, Google Search, Google Cloud services |
The paper's stated reason for the split is worth quoting in paraphrase: Gemini's reasoning is already good when all the relevant data is handed to it in natural language. Where it needs help is access to specialized data, special handling of non-language data such as geometry, and orchestration when many tools have to be sequenced and their outputs joined. Experts exist for the domains where the second problem is worst.
The operational framework is a closed cycle with three stations. Think & Plan: decompose the query into sub-tasks and pick a capability for each. Act: data operations, model inference, or model training, whichever the sub-task needs. Reflect & Recover: check the result; if a tool failed or returned nothing usable, revise the plan and go around again. The cycle runs until a final, grounded response exists, and the response carries its provenance: which data source each number came from.
The paper's worked example, as the agent runs it. The prompt: "As of September 23, 2024 12:00 UTC, what did Google's cyclone model say the predicted path of Hurricane Helene was? Get the list of counties in Florida with population > 20,000, and filter the list to the ones predicted to experience hurricane-force winds." Step through the plan; the map on the left updates with each tool's output and the loop on the right shows which station is active.
Trace the data types through those steps, because they are the whole design in miniature. Step 1 returns an ensemble: fifty tracks with wind radii. Step 2 collapses it to one polygon, the union of hurricane-force wind bands. Step 3 returns a table from the Census via Data Commons: county name, population. Step 4 returns polygons from the Places toolset, one per county. Step 5 is pure geometry: intersect polygons with the wind band, filter by population, and four counties survive. Step 6 puts them on the map in green and writes the sentence. No single model did anything hard. The agent's contribution was knowing that a cyclone forecast, a census table and a boundary service are three calls that compose.
Look at the agent's answer in the paper's Matanuska-Susitna flood example (Appendix A.3.3): a table of three zip codes with income and share over 60, then a numbered reasoning trace (flood forecast fetched; zip codes listed; polygons fetched; spatial intersection performed; statistics retrieved per zip), then a Data Provenance block with a source per column. That structure is not decoration. It is what the rubric-based grader in Chapter 10 checks for, and it is what makes an analyst trust the list enough to act on it. Base Gemini's answer to the same prompt names no source for its flood data, admits some percentages "were estimated," and lists eight zip codes where the authoritative flood polygon touched three.
# the Helene query as the agent's plan (schematic of ADK tool calls) fc = weather_expert.cyclone_forecast(storm="Helene", as_of="2024-09-23T12:00Z") # 50 members band = union(buffer(m.track, m.r_hurricane_km) for m in fc.members) # Polygon pop = demographics_expert.query("population of each county in Florida") # table via Data Commons geom = places.polygons([c.name for c in pop if c.population > 20000]) # {name: Polygon} hit = [n for n, g in geom.items() if g.intersects(band)] # 4 counties answer(table=hit, map=[band, *geom.values()], provenance={"forecast": fc.source, "population": pop.source})
How do you grade an agent that returns tables, maps and paragraphs? The paper builds two benchmarks with two different grading philosophies, and the second one is the more interesting because its limits are stated so plainly.
A hundred questions across four domains (Places 30, People & Communities 40, Weather & Environment 20, Multiple Domains 10), half descriptive-and-retrieval and half analytical-and-relational, every one answerable from public data and every one with a manually verified gold answer. Every question is phrased as fill-in-the-blanks so the grader can ignore prose: "Respond by filling in the blanks: There are __ botanical gardens in Lincoln, Nebraska." A subject-matter expert reviewed each question for being unambiguous, time-invariant in the short term, and useful to a real analyst.
Scoring: text blanks get ROUGE-L F1 against the gold. Numeric blanks get a clamped absolute percentage error:
The 10% dead zone is deliberate: businesses open and close, statistics get revised, two sources disagree at the margin, and none of that should cost points. The clamp at 100% error is a floor: being off by a factor of two is as wrong as being off by a factor of ten. A question's score is the mean over its blanks; the benchmark score is the mean over questions, averaged across five runs with a 95% confidence interval.
Left: the numeric scorer. Drag the prediction against a fixed gold value and read the score off the curve; note the flat top and the hard floor. Right: the paper's Table 11, drawn as bars. Toggle the category to see where the gap between the agent and plain Gemini opens up, and by how much.
| Q&A benchmark, mean of 5 runs | Places (30) | People & Communities (40) | Weather & Env. (20) | Multiple (10) | Overall (100) |
|---|---|---|---|---|---|
| Gemini 2.5 Flash + built-in tools | 0.35 | 0.46 | 0.35 | 0.28 | 0.39 ± 0.02 |
| Gemini 2.5 Pro + built-in tools | 0.45 | 0.54 | 0.54 | 0.38 | 0.50 ± 0.01 |
| Geospatial Reasoning Agent | 0.73 | 0.91 | 0.71 | 0.98 | 0.82 ± 0.02 |
The baselines are not straw men. Both Gemini agents are also built in ADK, with Google Search grounding, Google Maps grounding and code execution enabled, and are prompted to always give a best estimate and cite sources. They can look things up and run code; what they cannot do is call the models. Split by category, the agent beats Pro by 37% on descriptive questions (0.91 versus 0.67) and by 124% on analytical ones (0.74 versus 0.33). The harder the join, the wider the gap, which is the same pattern the Chapter 0 widget predicted.
The crisis prompts (Chapter 9's Helene query is a cousin of prompt 4) cannot have a gold answer, because they involve forecasts that change with every run and free-form outputs. So each prompt gets a rubric: a list of criteria such as "authoritative flood data was used," "there was a process to precisely identify which zip codes to include," "the response is an answer, not a plan for how someone could answer." An autorater built on Gemini 2.5 Flash reads the response against the rubric, writes a numbered justification per criterion, and issues a five-point verdict (not at all, somewhat, moderately, mostly, completely) mapped to 0, 0.25, 0.5, 0.75, 1.0. Each prompt runs ten times for each system.
| # | Crisis | Capabilities tested | Gemini 2.5 Pro | Agent |
|---|---|---|---|---|
| 1 | Flood | Flood forecast + Places | 0.18 | 0.80 |
| 2 | Flood | Flood forecast + Demographics | 0.28 | 0.80 |
| 3 | Typhoon | Cyclone forecast + Location | 0.23 | 0.93 |
| 4 | Hurricane | Cyclone forecast + Demographics | 0.45 | 0.80 |
| 5 | Drought | Weather forecast + Earth Engine | 0.58 | 0.88 |
| 6 | Extreme heat | Weather forecast + Demographics | 0.76 | 1.00 |
| 7 | Fire | Remote Sensing + Places | 0.40 | 0.84 |
| 8 | Flood | Remote Sensing + Demographics | 0.35 | 0.88 |
| 9 | Hurricane | TimesFM + Demographics | 0.00* | 1.00 |
| 10 | Hurricane | PDFM + Demographics | 0.25 | 0.91 |
| Mean (excluding 9) | 0.38 ± 0.17 | 0.87 ± 0.14 |
Row 9 is starred because base Gemini "was consistently unable to find authoritative labor statistics for 2024," so it could not attempt the question; the authors exclude it from the mean rather than count a free zero. Notice also that the base model is not hopeless: on extreme heat (weather forecast plus demographics, both of which search can partly reach) it scores 0.76. The gap is widest exactly where a polygon or a model call is unavoidable.
The paper devotes a full subsection to the limits of its own grading, and it is the most reviewer-like passage in the document. Two problems.
Consistent reasoning, inconsistent outputs. On the Maui fire prompt the agent's process was stable across ten runs, but the list of damaged schools and stores changed whenever the remote-sensing model flagged a different set of burned tiles. A rubric grades the process, so it cannot catch that the answer is unstable; only a gold answer could, and there is none.
Verifying methodology. For a widely covered event, a response might say it analyzed satellite tiles while actually synthesizing news reports found by search. The rubric checks for a claim of authoritative data use; it cannot easily check that the claim is true. The authors found this "particularly evident with high-profile events," where prompt refinement revealed that an apparently first-hand analysis was secondary sourcing.
A paper this broad earns trust by saying where it stops. Each of the four sections has a stated limitation, and each one is a research direction in disguise.
| Component | Stated limitation | What it implies |
|---|---|---|
| Remote Sensing Foundations | High-resolution RGB only; no temporal downstream tasks evaluated; oblique, multispectral and hyperspectral imagery not yet supported | Change detection over time (the thing crisis response most wants) is not benchmarked; a Sentinel-2 multispectral user still needs other backbones |
| Population Dynamics Foundations | Temporal look-back is a few months; more signals and finer granularity wanted; interpolation tests hide regions at random, not in the systematic gaps real outages create | The 0.85 R² is an optimistic bound for the case where a whole region's data goes dark at once |
| Model synergy | Aggregating 10 m AlphaEarth embeddings to tracts by histogram is crude; the fused model needs training per target variable; no single unified "meta-Earth" model exists yet | The obvious next paper: one model trained jointly on imagery, environment and population signals with a shared representation, replacing late fusion |
| Reasoning agent | Benchmarks test the agent's designed capabilities, not out-of-distribution queries or tool-orchestration failures; rubric grading has the two blind spots of Chapter 10 | The 0.82 and 0.87 are in-distribution numbers; the hard cases are exactly the ones not measured |
| Thing | Definition in one line | The number to remember |
|---|---|---|
| RS-SigLIP2 / RS-MaMMUT | 400M-param contrastive VLMs fine-tuned on RS-Landmarks (18M Gemini captions), RS-WebLI (3M mined), RS-Global (30M native-res) | RESISC45 zero-shot 72.40 → 80.13; RSITMD retrieval I2T 24.62 → 43.14 |
| RS-OWL-ViT-v2 + FLAME | Open-vocabulary detector, 2-stage cooldown fine-tune; FLAME picks uncertain-then-diverse examples for few-shot | DOTA mAP 13.77 → 31.83 zero-shot → 53.96 with 30 labels |
| RS-Global MTP backbone | MAE on 300M images, then multi-task (cls/seg/det) with alternating gradient descent | +14.93% avg over ImageNet ViT-L; FMoW 81.70, FLAIR 65.72, DIOR AP50 85.50 |
| Population Dynamics Foundations | GNN over regions; Maps + Knowledge-Graph search entities + busyness + weather → one vector per postal code; 17 countries; monthly | R² 0.85 across 17 countries; France from 7 neighbors: 0.52 GDP vs 0.00 for IDW |
| Environment models | Weather (hourly to 240 h), floods (severity + probability + polygon), cyclones (50 members, 15 days) | 50 scenarios, not one track |
| Synergy recipe | AlphaEarth 64 bands × 10-bin histograms (640) + area-weighted PDFM → GBDT, county-level split | FEMA R² 0.54 / 0.54 / 0.60 (+11%); CDC 0.42 / 0.56 / 0.60 |
| Forecast + covariates | TimesFM 2.0 zero-shot, then weather (dynamic) and PDFM (static) covariates | Cholera: −35.4% RMSE vs Prophet, a further −32.7% with weather; Ian: 2,575 predicted vs 2,496 damaged |
| Geospatial Reasoning Agent | ADK + Gemini 2.5; five experts, two toolsets; Think & Plan → Act → Reflect & Recover | Q&A 0.82 vs 0.50; crisis rubric 0.87 vs 0.38 |
| Scoring | ROUGE-L for text; 1.0 within ±10%, 0 beyond 100%, else 1 − APE; rubric → Likert → {0, .25, .5, .75, 1} | Five runs, 95% CI reported |
The pieces of Earth AI are each taught from zero elsewhere here. The contrastive image–text machinery of Chapter 2 is Contrastive Learning & CLIP, and the sigmoid-loss variant this paper fine-tunes has its own Veanor in EVA-CLIP's neighborhood. Masked autoencoding, the Stage 1 backbone objective, is worked through in Self-Supervised Learning and, for the audio cousin, in the Audio-MAE Veanor. The graph neural network that produces population embeddings is Graph Neural Networks plus the CS224W sequence starting at GNN I. FLAME's second move is k-means. Detection and segmentation heads are in Detection & Segmentation. The agent loop, tool use, and how to evaluate an agent with rubrics are covered by Agents & Tool Use, Agent Evaluation and the agent-evaluation survey Veanor. Knowledge distillation, the reason a small backbone can inherit a big model's competence, is Knowledge Distillation.
The synergy experiment is reproducible with public pieces. AlphaEarth embeddings are in Earth Engine; the US Population Dynamics embeddings and code recipes are on GitHub at google-research/population-dynamics; FEMA's National Risk Index and the CDC PLACES tract-level indicators are open downloads. Build the 640-feature histogram block, area-weight the zip vectors to tracts, split by county, and fit one boosted-tree model per target. If you get an 11% lift, you have reproduced the central claim of a fifty-author paper on a laptop. If you do not, you have found something worth writing up.