Aaron Bell, Amit Aides, Amr Helmy, Arbaaz Muslim, Aviad Barzilai, Aviv Slobodkin, Bolous Jaber … Niv Efron, Shravya Shetty (Google Research · Google X · Google Cloud · Google Geo) — arXiv:2510.18318, October 2025 (v4, February 2026)

Distill the Planet, Then Reason

Satellite pixels, search trends and storm forecasts live at different resolutions, on different clocks, in different silos. Earth AI turns each silo into a foundation model, shows that fusing their embeddings lifts almost every prediction it tries, and puts a Gemini agent on top that plans, calls the models, and answers crisis questions an analyst would spend a day joining by hand.

Prerequisites: what an embedding is (a fixed-length vector that stands in for a thing) + what a contrastive image–text model does (CLIP-style: matching pictures to captions). Foundation models, graph networks, ensembles, boosted trees and agents are all built from zero here.
12
Chapters
14
Simulations
+11%
R² from fusing two embeddings
0.82 vs 0.50
Agent vs plain Gemini 2.5 Pro

Chapter 0: The Problem

It is the evening of September 23rd, 2024. A hurricane named Helene is spinning up in the Gulf of Mexico, and an emergency manager in Florida needs one list before she goes home: every county with more than 20,000 people that sits inside the forecast band of hurricane-force winds. Not the whole state. Not a news article. A list with populations next to it, sorted, so that cooling centers, buses and water can be staged tomorrow morning.

Every piece of that answer exists somewhere. The cyclone forecast lives in a weather lab as fifty possible storm tracks. County boundaries live in a mapping service as polygons. Populations live in the Census Bureau, mirrored through a statistics service. None of those three systems knows the other two exist. So the analyst downloads the forecast, exports the polygons, pulls the census table, converts three coordinate systems into one, computes which polygons overlap which wind band, joins the population column, filters, sorts. If she is fast, that is an afternoon. Helene made landfall three days later.

That afternoon is the problem this paper is about. Geospatial data is not scarce. It is siloed: it comes in three families that were never designed to be read together.

Three Silos, Three Resolutions, Three Clocks

Each lane below is one family of Earth data drawn at its native grain: the size of one cell is the size of one measurement. Pick a question and watch which lanes it needs and how many hand-made joins connect them. The counter is the analyst's afternoon.

The paper sorts questions like these into a hierarchy of three rungs, and the widget follows it. A descriptive question ("the hottest day in New York in July 2024") touches one silo and needs no join: look it up. An analytical question ("how many hospitals were in the storm band when Katrina came ashore") touches two silos and needs a spatial join: a polygon from one system intersected with points from another. A predictive question ("which Indian cities will have the most vulnerable people exposed to flooding by November 2027") touches all three and needs a join plus a model: nobody has that number in a table, it has to be forecast.

Why the Silos Never Merged

It is tempting to think the fix is one giant table. Put every pixel, every postal code and every forecast in one database and query it. The widget shows why that fails before it starts: the three families do not share a grain.

Imagery is measured per pixel, and a pixel can be anything from ten centimeters to ten meters wide depending on the satellite. A single city block is a few hundred pixels at one resolution and a few hundred thousand at another. Population is measured per administrative region: a postal code, a census tract, a county. Those are irregular polygons that were drawn for mail delivery and elections, not for science, and they change size by a factor of a thousand between Manhattan and Montana. Environment is measured on a regular grid with a clock attached: a forecast is a value per grid cell per hour, and it is reissued every few hours, so the same location has many competing values dated by when they were computed.

The silos are not a data-engineering accident. They are three different answers to "what is one measurement?" A pixel, a polygon and a grid-cell-hour are three incompatible units. You cannot average a postal code with a satellite tile any more than you can add meters to seconds. Any system that wants to reason across them has to first translate each into something comparable, and that translation is exactly what a foundation model's embedding is for.

There is a second, quieter reason the silos never merged: sparsity. Labels are precious. The census reports household income per tract once every few years; a river gauge exists on some rivers and not others; a human-annotated satellite image of a flooded road is a rare thing. Any method that needs a large labeled dataset per question is dead on arrival. Whatever lives above the silos must be able to learn a new target variable from a few hundred labeled regions, not a few million.

What the Field Tried Before

Geospatial AI did not start with agents. It started with specialized models: one network for land cover, one for buildings, one for crop type, each trained on its own dataset and speaking to nobody. The next step was general-purpose foundation models for Earth observation, such as Prithvi and SatMAE, pretrained on large piles of multispectral satellite imagery so that one backbone could be fine-tuned for many tasks. That solved half of the imagery silo. It did nothing for population or environment, and it produced a backbone, not an analyst.

Most recently, the field started pointing large language models at geospatial data and building benchmarks to test them: MapEval for map reasoning, GeoChain for chain-of-thought geography. Those experiments exposed a gap that this paper is built around. A language model reasons well over natural-language facts it has been handed. It reasons badly over a polygon, a raster, or a table of forecasts it has to go find, parse and join itself.

Era 1 · Specialized models
One network per task (buildings, land cover, roads). Accurate, mute, and blind to the other silos.
↓
Era 2 · Earth-observation foundation models
One imagery backbone, many fine-tunes. Solves the imagery silo's label scarcity; still one silo.
↓
Era 3 · This paper
A foundation model per silo, evidence that their embeddings are complementary, and an agent that plans the joins.

What "Complementary" Will Mean

Hold onto one number for the rest of the lesson. When the authors predict FEMA's twenty disaster-risk indices for every census tract in the country, using landscape embeddings alone gets an average R² of 0.54. Using population embeddings alone also gets 0.54. Using both gets 0.60. The two silos know different things about the same tract, and the whole paper is an argument that this is true everywhere: imagery knows the physical hazard, population knows who can absorb it, environment knows when it arrives.

The analyst wants "counties in Florida with more than 20,000 people inside Helene's hurricane-force wind band." Why is this an analytical question rather than a descriptive one, in the paper's hierarchy?

Chapter 1: The Key Insight

The paper's own one-sentence summary is buried at the end of its architecture section, and it is the sentence to remember: capable models distill reality, and a powerful agent reasons over that distillation. Two layers, with a deliberate division of labor between them.

The bottom layer is three families of foundation models, one per silo. Each takes its silo's raw, incompatible unit (pixels, polygons, grid-cell-hours) and produces something a downstream system can actually use: a fixed-length embedding vector, a detection, a classification, or a forecast with uncertainty attached. That is the distillation. Reality goes in at whatever grain it has; a compact, comparable representation comes out.

The top layer is a single Gemini-powered agent that never touches a pixel. It reads the question, decomposes it into steps, decides which model or tool each step needs, calls them, checks the result, and writes the answer. That is the reasoning. The agent's competence does not come from knowing geography; it comes from having a menu of specialists and knowing how to sequence them.

Earth AI, Component by Component

This is the paper's Figure 1 rebuilt so you can interrogate it. Click any box to see what goes in, what comes out, and what it is trained on. Watch the data types change across the diagram: raw pixels become vectors, vectors become numbers, numbers become a sentence with a map attached.

Click a component.

Read the diagram left to right and notice that the arrows carry different kinds of things. The leftmost arrows carry raw data: a stack of satellite tiles, a table of Search Trends per postal code, a gridded weather forecast. The middle arrows carry distillations: a 64-number embedding per ten-meter pixel from AlphaEarth, a fixed-length vector per postal code from Population Dynamics Foundations, fifty candidate storm tracks from the cyclone model. The rightmost arrows carry arguments: a plan, a tool call, a table, a sentence. The layers are separated by the type of thing that flows between them, not by an org chart.

Why Split It This Way

Why not train one enormous model on everything and ask it questions directly? The paper gives three reasons, and each one is visible in the diagram.

First, label scarcity lives at the bottom. A foundation model per silo can be pretrained once, self-supervised, on the ocean of unlabeled data each silo has (300 million satellite images; every postal code's search behavior), and then adapted to a new target variable with a few hundred labels. If the agent had to learn geography from question–answer pairs, it would need labels that do not exist.

Second, complementarity lives in the middle. Because each silo's model produces an embedding in its own space, an analyst (or the agent) can concatenate two of them and train a tiny model on top. Chapter 7 is entirely about this: the same boosted-tree recipe, fed landscape embeddings plus population embeddings, beats either alone on 39 of 41 target variables.

Third, extensibility lives at the top. The agent's capabilities are organized by domain, as tools or "expert" sub-agents. Adding a new model means adding a new tool; the reasoning layer does not retrain. The authors say this explicitly: the domain-oriented design exists so they can plug in new geospatial domains and swap implementations as language models and frameworks improve.

The insight is not "use an LLM for geography." It is "put a reasoning layer above a distillation layer, and make the seam between them a type: embeddings, detections and forecasts go up; plans and tool calls go down." A base Gemini with web search sits on the wrong side of that seam. It can read news about a flood; it cannot intersect a flood polygon with postal codes, because nobody handed it a polygon. Chapter 10 measures exactly how much that costs: 0.50 versus 0.82 on the same hundred questions.

The Three Findings, and Where Each Lives in This Lesson

FindingEvidence in the paperChapters
State-of-the-art models per siloRemote Sensing Foundations: best zero-shot classification and retrieval on most public benchmarks; 31.83% mAP open-vocabulary detection on DOTA versus 13.77% for the base model. Population Dynamics Foundations: R² 0.85 across 17 countries; independently validated by seven outside groups.2, 3, 4, 5, 6
Synergy across silosFusing AlphaEarth and Population Dynamics embeddings lifts FEMA-risk R² by 11% on average and beats either alone on all 21 CDC health variables; weather covariates cut cholera forecast error by a third; a cyclone ensemble plus population embeddings predicted Hurricane Ian's damaged-building count within 3% three days out.7, 8
Agentic reasoningThe Geospatial Reasoning Agent scores 0.82 on a 100-question fill-in-the-blank benchmark versus 0.50 for Gemini 2.5 Pro with search, maps and code tools; 0.87 versus 0.38 on ten rubric-graded crisis scenarios.9, 10

How the Idea Was Probably Found

The paper reads like the third step of a sequence you can reconstruct. Google had already shipped the pieces separately: a building-detection model, a flood-forecasting system, a population-embedding model released in 2024, a cyclone model in 2025, AlphaEarth in 2025. Each team benchmarked its own model on its own tasks. The first question someone asked was probably the cheap one: "if I have the population embedding and the AlphaEarth embedding for the same tract, does stacking them help?" It did, by 11%. The second question follows immediately: "if the models are complementary, who is going to do the stacking for a user who is not a data scientist?" The answer to that is an agent, and the benchmark of ten crisis scenarios is what you build when you want to prove the agent can do the stacking on its own.

The agent in Earth AI never sees a satellite pixel. Which statement best explains why that is a design choice rather than a limitation?

Chapter 2: Seeing in Words

Start with the imagery silo, because it is the one where the paper's models are most concrete and its benchmarks most public. The first capability the authors want is simple to state and was, until recently, impossible: type "a solar farm next to farmland" and have a model find every satellite tile that matches, with no training examples of solar farms at all.

The way to get there is a vision–language model in the CLIP family: two encoders, one for images and one for text, trained so that an image and its caption land close together in a shared embedding space and an image and a stranger's caption land far apart. Once that space exists, classification is free: embed the tile, embed a handful of candidate labels as sentences, and pick the label whose vector is closest. Retrieval is the same trick run the other way.

The Joint Embedding Lab

Six synthetic overhead tiles on the left, six text prompts across the top, and the cosine similarity between every pair in the grid. The base encoder was trained on ordinary web photos, so it half-knows what a runway looks like from above. Toggle the remote-sensing fine-tune and watch the diagonal sharpen: that is what 18 million caption pairs of overhead imagery buy. Click a prompt to run retrieval; click a tile to run zero-shot classification.

The widget is a cartoon of the real thing, but the cartoon is faithful in one respect: a general-purpose model is not wrong about overhead imagery, it is blurry. Off-the-shelf SigLIP2 gets 72.40% top-1 zero-shot accuracy on RESISC45, a 45-class scene dataset. That is far above chance. The remote-sensing version of the same architecture, trained on the same recipe plus the datasets below, gets 80.13%. On the harder FMoW benchmark the jump is 41.25% to 48.13%. Nothing about the architecture changed. Only the captions did.

Where 18 Million Overhead Captions Come From

The bottleneck for a remote-sensing VLM is not images. Google has decades of aerial and satellite imagery. The bottleneck is text: nobody writes captions for satellite tiles. The paper's answer is three datasets, each built by a different trick.

RS-Landmarks · 18M images, Gemini-written captions
Take an aerial or satellite image. Look up every place and landmark from Google Maps whose footprint falls inside it (a stadium, a marina, a school). Hand the image and that list to Gemini 1.5 Pro with a curated prompt, and ask for a concise caption. The Maps data grounds the caption in names the model could never guess from pixels alone.
↓
RS-WebLI · 3M clean images mined from web image–text pairs
WebLI is a giant web image–caption corpus that happens to contain some overhead shots. Finding them is a bootstrapping ladder: hand-label a few hundred, train a classifier, get 40K candidates, crowd-label 60K with random negatives, train better classifiers, run them over all of WebLI, keep 3 million. Then repeat on the 100-billion-image WebLI to get a much larger variant.
↓
RS-Global · 30M images at native resolution, every land area
Aerial and satellite images from 10 cm to 10 m per pixel, 2003 to 2022 weighted toward newer, covering all land except the poles and remote islands, biased toward where humans are. Each image gets a Gemini 1.5 Pro description. This is the balance dataset: Landmarks is rich but urban, Global is everywhere.

Two model families are fine-tuned on those datasets: MaMMUT (a joint encoder–decoder design) on RS-Landmarks plus the smaller RS-WebLI, and SigLIP2 (a sigmoid-loss contrastive model) on the larger RS-WebLI plus RS-Landmarks plus RS-Global. Both use 400-million-parameter image encoders and 400-million-parameter text encoders, sized so that the models can run on constrained hardware, not just in a datacenter.

Think of the three datasets as three different teachers. RS-Landmarks teaches the model the names of things (this blob is "Dodger Stadium"). RS-WebLI teaches it how people describe overhead views (whatever the web caption said). RS-Global teaches it what the planet looks like at every resolution, so that a tile of Nigerian farmland is as familiar as a tile of Ohio.

The Benchmarks, With the Numbers

Zero-shot classification asks: given a tile and the names of the 45 (or 30, or 21) classes as text, does the nearest class embedding match the ground truth? Retrieval asks the symmetric questions: given a caption, is the right image among the top results (text-to-image), and given an image, is the right caption (image-to-text)? The paper reports the average of top-1, top-5 and top-10 for retrieval.

Zero-shot top-1 accuracyFMoWSkyScriptRESISC45UCMAID
SigLIP2 400M (base)41.2565.9472.4081.4375.56
RS-SigLIP2 400M48.1368.1380.1384.8678.26
RS-MaMMUT 400M47.2469.4672.3180.2971.96
RemoteCLIP (prior RS model)––79.84–91.30
Retrieval, avg of top-1/5/10RSICD I2T / T2IUCM-Captions I2T / T2IRSITMD I2T / T2INWPU I2T / T2I
SigLIP2 400M (base)27.36 / 27.9370.48 / 68.5224.62 / 25.1431.86 / 36.70
RS-SigLIP2 400M38.37 / 37.6476.67 / 75.3343.14 / 47.2645.12 / 37.74
RS-MaMMUT 400M33.33 / 33.5974.76 / 71.7942.63 / 42.5841.44 / 32.28

Read the tables honestly, the way a reviewer would. On four of five classification benchmarks the remote-sensing fine-tune beats its own base model, and on three of them it beats every prior remote-sensing model. But RemoteCLIP still wins AID by a wide margin (91.30 versus 78.26), and the SkyScript number carries an asterisk because that model was evaluated in-domain. Retrieval is the cleaner story: RS-SigLIP2 wins seven of eight columns, and the gain on RSITMD (24.62 to 43.14 image-to-text) is close to a doubling. The paper's summary phrase is "state of the art on the vast majority of benchmarks," and the tables support exactly that phrase, no stronger.

What Is Actually Being Computed

A tile x of shape [3, 224, 224] goes through the image encoder and comes out as a single vector v of a few hundred to a thousand numbers, then normalized to unit length. A prompt string goes through the text encoder and comes out as a unit vector u. The score is their dot product, which for unit vectors is the cosine of the angle between them:

score(x, prompt) = v · u = cos θ

Zero-shot classification over K class names is K dot products and an argmax. Retrieval over a million tiles is a million dot products, which is why the whole design cares about pre-computing and caching image embeddings: encode every tile once, then any new question is a matrix multiply against the cache. That same caching idea reappears in the detection model of the next chapter and in the agent's imagery toolset in Chapter 9.

# zero-shot classification with a contrastive VLM (shapes in comments)
v = normalize(image_encoder(tile))              # [D]      one unit vector per tile
U = normalize(text_encoder(["a photo of a %s" % c for c in classes]))  # [K, D]
scores = U @ v                                     # [K]      K cosines
label = classes[scores.argmax()]

# retrieval: same math, cached the other way round
V = normalize(image_encoder(all_tiles))          # [N, D]   computed ONCE, stored
u = normalize(text_encoder("a solar farm next to farmland"))
top = (V @ u).topk(10)                            # [10]     a million dots, no re-encoding
The engineering decision worth copying: the models were kept at 400M parameters on purpose, and the paper notes that with 7-billion-parameter chat-style remote-sensing models such as GeoChat and LHRS-Bot they are "comparable." A 17× smaller model that can be batched over a million cached tiles is worth more to an agent than a chatty giant that has to look at each image one at a time.
The RS-SigLIP2 model uses exactly the same architecture and loss as base SigLIP2. What is the single most important reason it beats the base model on remote-sensing benchmarks?

Chapter 3: Naming the Unseen

Classifying a whole tile is a coarse tool. The crisis-response prompts in this paper ask things like "which schools and grocery stores in Lahaina were in the burned area" and "how many lakes are in the residential parts of this zip code." Those need boxes, not labels: where, in the image, is each object of a kind I name in words. That is open-vocabulary object detection (OVD), and cataloging every possible class up front is hopeless for the planet, so the vocabulary has to be open.

The base detector is OWL-ViT v2, the same idea as the VLM but with a twist: the image encoder produces one embedding per candidate box instead of one per image, and the text encoder produces one embedding per query. A box is a detection of "grain silo" if its embedding is close to the embedding of the words "grain silo." Because the two encoders are independent, the box embeddings for a region can be computed once and cached, and any new query is a dot product against the cache. The paper calls this split processing and it is the reason the agent can re-query the same satellite tiles cheaply.

Training the Detector on Overhead Imagery

Zero-shot OWL-ViT v2 gets 13.77% mean average precision on DOTA and 14.98% on DIOR, two standard aerial detection benchmarks. Overhead objects are small, oddly oriented and packed together; a detector trained on web photos of dogs and bicycles half-recognizes them. The remote-sensing version, RS-OWL-ViT-v2, more than doubles both numbers to 31.83% and 29.39%, with a two-stage cooldown fine-tune:

Stage 1 · 2 epochs on RS-WebLI with pseudo-boxes
The 3M web overhead images have captions but no boxes. Run the original OWL-ViT over each image using its alt text (or a fixed label list) as queries, keep only boxes above a confidence threshold, and treat those as training labels. Noisy, but plentiful.
↓
Stage 2 · 32 epochs on 67,000 human-annotated aerial images
An internal dataset with 3.5 million labeled instances across 34 categories. Clean, small, expensive. The learning rate decays linearly through both stages, so the noisy data shapes the model first and the clean data finishes it.

Why that order? Because the noisy pseudo-boxes teach what overhead objects look like in general (a thing with a shadow, a rectangle at a strange angle), while the human labels teach precision. Running them the other way round would let 3M noisy images wash out the precise ones.

When Words Are Not Enough: FLAME

Even a good zero-shot detector fails on ambiguity. "Silo" could mean a grain silo or a missile silo; "boat" at ten-meter resolution is a few pixels. The classic fix is few-shot learning: give the detector a handful of labeled examples of the thing you mean. The expensive part is choosing those examples: a human has to annotate them, so you want the ones that teach the most per click.

FLAME (from the same group) is a one-step active-learning strategy for exactly that. It runs the zero-shot detector, looks at the embeddings of all the candidate boxes it produced for your query, and picks the examples to show a human in two moves.

FLAME: Uncertainty, Then Diversity

Each dot is one candidate box the detector found for a query, drawn in a 2-D projection of its embedding. Dense clumps are boxes the model is confident about (many look alike). Step 1 keeps the low-density margins: the ambiguous ones. Step 2 clusters those with k-means and takes the point nearest each center, so the human labels examples that are both uncertain and different from each other. Slide the budget to see how few labels FLAME needs; the mAP readout follows the paper's numbers.

Step 1, uncertainty filtering. Combine each box's image features with the text query into one embedding, reduce its dimensionality, and estimate the density of the resulting cloud. Points in dense regions are boxes the model has seen many near-copies of; it is sure about them. Points on the low-density margins of the core are the ambiguous ones. Keep the margins.

Step 2, diversity sampling. The margins may still contain fifty near-duplicate ambiguous boxes. Run k-means on them with k equal to your labeling budget and take, from each cluster, the sample closest to its center. Now the human sees k examples that are uncertain and spread out.

With 30 labeled examples per category, this lifts mAP on DOTA from 31.83% (zero-shot) to 53.96%, and on DIOR from 29.39% to 53.21%. That beats SIoU, a recent few-shot detection method built specifically for small objects, on both benchmarks (45.88% and 52.85%). Thirty clicks per class, chosen well, nearly doubles the detector.

The lesson under the lesson: a foundation model's zero-shot answer is a starting point, not a verdict. The agent in Chapter 9 uses the imagery toolset in exactly this spirit: a query returns candidates, and a cheap adaptation step, seeded by the model's own uncertainty, turns "roughly right" into "right." Choosing which thirty examples is worth more than choosing thirty at random, because the model already knows which of its guesses it doubts.
# FLAME in ten lines (the paper's two moves)
E = embed_boxes(image_feats, text_query)      # [N, D]  one vector per candidate box
Z = reduce(E, dims=2)                         # [N, 2]  cheap to estimate density in
dens = kde(Z)                                # [N]     how crowded is each point's neighbourhood
margin = Z[dens < quantile(dens, 0.3)]        # the uncertain 30%
centers, assign = kmeans(margin, k=budget)
picks = [margin[assign == j][argmin(dist(margin[assign == j], centers[j]))] for j in range(budget)]
# show `picks` to a human; train the lightweight few-shot classifier on their labels
FLAME picks labeling candidates from the low-density regions of the box-embedding cloud, then k-means among them. Why the low-density regions specifically?

Chapter 4: The Backbone

The VLM and the detector both answer questions in words. Plenty of geospatial work does not: segmenting a flooded road pixel by pixel, outlining every building in a city, classifying scenes for a dataset that will be fine-tuned anyway. For those, what you want is a backbone: a vision transformer whose features are so good that a small task head on top, trained on a few thousand labels, beats a model trained from scratch on millions.

The paper's backbone recipe, called RS-Global MTP, has two stages, and the second one is the unusual part.

Stage 1: Masked Autoencoding on 300 Million Tiles

Take the image-only version of RS-Global and scale it ten times, to 300 million aerial and satellite images, randomly cropped and resized during training so that the effective resolution varies continuously from 10 cm to 10 m per pixel. Then train a masked autoencoder (MAE): chop each image into patches, hide most of them, and ask the network to reconstruct the hidden pixels from the visible ones. No labels anywhere.

Masked Autoencoding, Then Alternating Tasks

Left: an 8×8 patch grid of a synthetic overhead scene. Drag the mask ratio and press Reconstruct to watch the encoder fill in what it never saw: the only way it can is by having learned what fields, roads and roofs look like. Right: Stage 2's schedule. Multi-task pretraining does not sum three losses in one step; it alternates, one task per step, so the encoder hears one teacher at a time.

Why does hiding three quarters of an image teach anything? Because the reconstruction target is not a class name, it is the world itself. To fill in a hidden patch of farmland, the encoder must have learned that furrows continue in straight lines, that a road that enters a tile usually exits it, that a roof has a shadow on the side away from the sun. Those are the same regularities every downstream task needs. The paper credits MAE with capturing "low-level spectral patterns as well as higher-level spatial structures unique to remote sensing data."

Stage 2: Multi-Task Pretraining With Alternating Gradient Descent

MAE features are general but not task-shaped. So the second stage adds three supervised objectives on internal labeled data: scene classification, semantic segmentation, and object detection. The naive way to combine them is to compute all three losses on each batch, add them up, and take one gradient step on the sum. The paper does something else: Alternating Gradient Descent (AGD), where each training step optimizes one task, cycling through them.

naive: θ ← θ − η ∇(Lcls + Lseg + Ldet)
AGD:   θ ← θ − η ∇Lcls, then θ ← θ − η ∇Lseg, then θ ← θ − η ∇Ldet, repeat

The reason is task interference. Segmentation wants the encoder to keep fine boundaries; classification wants it to throw boundaries away and summarize; detection wants something in between. Summing the losses makes every step a compromise none of the three tasks asked for, and the gradients can partially cancel. Alternating lets each task's head and its share of the encoder specialize for one step while the shared representation still has to serve all three over time. The paper's phrase: it "mitigates task interference, allowing different task experts to specialize while still contributing to a unified representation."

A worked intuition: imagine three coaches training one athlete. Summing the losses is all three shouting corrections at once during every rep, and the athlete averages them into a motion nobody wanted. Alternating is one coach per rep. Over a session the athlete still learns all three skills, but each rep has a clear target. AGD is that, for gradients.

How Much It Buys

The backbone was evaluated by fine-tuning it on 13 downstream benchmarks across four task types, always with the encoder weights tunable, the same batch size and epochs for every model, and only the learning rate tuned per model. Results are reported relative to a ViT-L/16 pretrained on ImageNet, the standard off-the-shelf starting point:

Backbone (relative to ImageNet ViT-L/16)Semantic seg. (mIoU)Detection (AP50)Classification (acc.)Instance seg. (mAP)
Random init−16.23%−17.01%−24.14%−40.87%
DINOv2 ViT-L/14+4.44%+4.69%+4.89%+75.94%
SigLIP2 So400m+4.50%+2.47%+4.74%+63.75%
RS-SigLIP2 (the Chapter 2 model)+4.15%+2.52%+4.79%+66.74%
RS-Global MTP+6.09%+5.91%+4.58%+82.92%

Three things to notice. The instance-segmentation column is enormous for every pretrained backbone, which tells you the ImageNet baseline is simply bad at outlining buildings; the interesting comparison is RS-Global MTP against DINOv2, and the remote-sensing backbone wins by seven points there. On classification the remote-sensing backbone is slightly behind DINOv2 (4.58% versus 4.89%), which the paper does not hide. And the VLM from Chapter 2, when used as a backbone, is no better than base SigLIP2: contrastive image–text training makes a great retriever and an ordinary backbone. Different objectives, different strengths. Averaged over all task categories, RS-Global MTP improves 14.93% over the ImageNet baseline.

Against published state of the art on public splits the backbone also leads: 81.70% on FMoW classification (prior best 79.30%), 65.72% mIoU on FLAIR segmentation (63.10%), 85.50% AP50 on DIOR detection (81.10%).

# Stage 2 in code: alternate, do not sum
tasks = [("cls", cls_loader, cls_head), ("seg", seg_loader, seg_head), ("det", det_loader, det_head)]
for step in range(total_steps):
    name, loader, head = tasks[step % 3]        # one task per step, round robin
    x, y = next(loader)
    feats = encoder(x)                          # shared ViT, MAE-initialised
    loss = head.loss(feats, y)                  # ONLY this task's loss
    loss.backward(); opt.step(); opt.zero_grad()
In Stage 2 the three supervised losses are applied on alternating steps rather than summed into one loss. What problem is that choice solving?

Chapter 5: Population Dynamics

Now the second silo, and the one with no pixels in it at all. What does a postal code look like to a model? Not from above: from the inside. What businesses are there, what people search for, how busy the streets are at noon, what the weather was. The Population Dynamics Foundations model (PDFM) turns all of that into one vector per region, and that vector turns out to predict things nobody put into it: diabetes rates, home values, night-time lights, dengue.

What Goes In

Four families of signal, all aggregated to an administrative region (a US postal code, or a similarly sized region elsewhere), all privacy-preserving by construction because nothing below the region level ever exists in the model:

SignalSourceWhat it encodes
Built environmentGoogle MapsCounts and kinds of places: restaurants, clinics, schools, parks, road density
Human behaviorSearch TrendsWhich topics people in the region search for, relative to elsewhere
Human presenceAnonymized busynessHow occupied places are through the day and week
Natural environmentWeather and air qualityTemperature, precipitation, pollution over the period

The Search Trends channel is the clever one, and the global version of the model changed how it is built. Instead of counting the top raw queries per region (which are language-specific and privacy-sensitive), the global model counts the top Knowledge Graph entities that the queries resolve to. "Taylor Swift boyfriend" and "KC tight end" both map to the entity Travis Kelce. For each region, the model records the frequencies of the top 500 entities searched at least 20 times; entities are ranked by how many regions they appear in across 17 countries, and the top 1,000 form a fixed vocabulary. The least popular entity in that vocabulary still appeared in searches from more than 9,000 regions, which is the privacy floor: nothing rare enough to identify anyone survives.

Why entities beat queries, in one line: mapping many phrasings, in many languages, to one thing makes the signal comparable across countries and coarser about individuals at the same time. Generalization and privacy pull in the same direction here, which is rare.

How the Signals Become One Vector

Regions are not independent. A postal code shares a border, a river, a commute pattern with its neighbors. So the model is a graph neural network: each region is a node holding its four signal vectors, edges connect nearby regions, and a few rounds of message passing let each node's embedding absorb its neighborhood. The training objective is self-supervised (the original PDFM paper trains the GNN to reconstruct held-out signals from neighbors), so the embedding is a compressed summary of "what this region and its surroundings are like" that was never told about any downstream target.

A Region Graph, and Why Embeddings Beat Distance

A hex map of synthetic regions, each with four signal channels (Maps, Search, busyness, weather) shown as its color mix. Press Message pass to watch each node blend in its neighbors: that is the GNN, and the result is the embedding. Then Hide 20% of the regions' target value and compare two ways of guessing it: inverse-distance weighting (the classic GIS method, which only knows where a region is) versus nearest neighbors in embedding space (which knows what a region is like). The R² readouts are computed live on this toy; the paper's real numbers are in the table below.

The toy makes the paper's central claim about this silo visible. Inverse-distance weighting assumes that nearby regions are similar. Embedding-space neighbors assume that similar regions are similar, wherever they are: a university town in Ohio has more in common with a university town in Oregon than with the farmland next door. When the hidden target correlates with what a region is like rather than where it sits, the embedding wins.

The Two Expansions in This Paper

The 2024 PDFM covered the contiguous United States for a single year. This paper adds two things.

Global coverage, 17 countries. Australia, Belgium, Brazil, Canada, France, Germany, India, Italy, Japan, Mexico, the Netherlands, Nigeria, Portugal, Spain, Switzerland, the United Kingdom and the United States. The embeddings are built in one shared space, so a model trained on US labels can be applied in the UK or Brazil. The benchmark is spatial interpolation: hide 20% of a country's regions, predict four globally available targets from Earth Engine (elevation, night-time lights, population density, tree cover), score with R².

Global PDFM, R² (hold out 20% of regions)ElevationNight lightsPop. densityTree coverMean
United States0.980.930.880.870.92
India0.970.900.870.850.90
Japan0.910.940.940.810.90
Netherlands0.870.880.890.420.77
Nigeria0.790.520.520.710.63
All 17 countries0.910.880.880.730.85

Notice where it is weak, because the weaknesses are informative. Tree cover in the Netherlands (0.42) and Switzerland (0.50) is hard because those countries have little of it and what exists is not where the searches and businesses are. Nigeria is the lowest overall (0.63) with night lights and population density at 0.52: the input signals are sparser where Maps coverage and search volume are lower. The embedding is only as rich as the digital footprint it summarizes.

The stronger test is cross-country extrapolation: train on regions in Belgium, Switzerland, Germany, Spain, the UK, Italy and the Netherlands, predict France, using Eurostat's NUTS-3 statistics. Against the two standard GIS interpolators:

Predicting France from 7 neighborsGDP/capita R²MAPEDeath rate R²MAPEFertility R²MAPE
Inverse Distance Weighting0.0028.03%−0.0719.29%0.518.78%
Radial Basis Functions0.1325.71%0.5212.91%0.509.35%
PDFM embeddings0.5219.03%0.787.99%0.716.79%

Read the first row. Inverse-distance weighting gets an R² of exactly zero on GDP per capita, and a negative R² on death rate: it is worse than predicting the mean. That is what happens when you extrapolate into a country using only the coordinates of regions in other countries. The embedding does not know France is France; it knows a NUTS-3 region full of pharmacies, ski searches and low busyness in August looks like certain Swiss and Italian regions, and it borrows their numbers.

Temporal embeddings, monthly since July 2023. The static embedding is a snapshot. The temporal version rebuilds it every month from that month's signals only. Tested on monthly emergency-department visits for COVID-19, flu and RSV in a fixed holdout of ten US states over two years, the time-matched embedding consistently reduces mean absolute error compared with a static embedding frozen in July 2023, with the largest gains in fall and winter when disease burden is highest. That is the "Temporal drift" toggle in the widget: a static snapshot slowly goes stale as behavior shifts, and a monthly embedding tracks it.

Seven Outside Groups Checked

Foundation-model papers usually benchmark themselves. This one also reports independent validations of the static US embeddings by groups that had no stake in the result, and the table is worth its space because the tasks are so different from each other:

WhoTaskResult
CARTOInterpolate their proprietary Human Activity Index across the US, random forest on PDFMR² 0.882 on the test split
CARTOIowa liquor salesPDFM alone R² 0.666 (0.097 above demographics alone); ensemble with demographics 0.707
SustGlobalHome-insurance premium tier, XGBoost on PDFM93% top-2 accuracy at county level, 89.5% at zip level
University of Oxford1–12 month dengue forecasting in BrazilBest model R² 0.849 → 0.901 with PDFM; with TimesFM at 6 months 0.500 → 0.686, at 12 months 0.456 → 0.656
Institute for Disease ModelingPoliovirus risk in Nigeria, ridge regression with spatial effectsBetter fit (DIC/WAIC); predicted risk changed more than 2-fold in 18% of local government areas
Mount Sinai / Boston Children's / HarvardMeasles vulnerability mapsSuper-resolved from county level to zip-code level
The dengue row is the one to remember for Chapter 8. PDFM helped most when paired with TimesFM at 6- and 12-month horizons, lifting R² from 0.50 to 0.69 and from 0.46 to 0.66. A static "what is this region like" vector is exactly the covariate a time-series model lacks, and that pairing is the recipe the cholera forecast uses.
Inverse-distance weighting scores R² = 0.00 on French GDP per capita when trained on seven neighboring countries, while PDFM embeddings score 0.52. What does the embedding know that coordinates do not?

Chapter 6: Environment

The third silo is the one with a clock. Imagery and population describe the world as it is or was. Environment models describe what it will be, with uncertainty attached, and they are reissued constantly. The paper does not train new environment models; it integrates three that Google already runs, and what matters for the rest of the lesson is the shape of what each one returns.

ModelWhat it returnsHorizon and history
Weather (Maps Platform Weather API, MetNet-family nowcasting)Hourly temperature, precipitation, wind, UV per locationHourly to 240 hours, daily to 10 days
Floods (Flood Forecasting API)Riverine flood predictions from gauge data: inundation area, severity level, probabilityCurrent, plus history back to August 1st, 2025
Cyclones (experimental, stochastic neural networks)Formation, track, intensity, size and shape as 50 possible scenariosUp to 15 days ahead; history back to January 1st, 2022

The cyclone model deserves the widget, because "50 possible scenarios" is the most important data type in this paper's crisis chapters and the least intuitive one.

Fifty Futures for One Storm

A storm starts at the dot on the left. Each thin line is one ensemble member: a plausible track sampled from the model. Slide the lead time and watch the fan open. The shaded band is the union of every member's hurricane-force wind radius, and the grid behind it is a state of counties with populations. The readout counts counties over 20,000 people that any member's wind band touches: the exact computation the agent performs for Hurricane Helene in Chapter 9.

Two things the widget teaches. First, an ensemble is not a single answer with error bars; it is fifty complete, internally consistent futures. Any question you ask of the storm ("which counties get hurricane-force winds") can be asked of each member separately, which is how the damage model in Chapter 8 produces a distribution of damaged buildings rather than one number. Second, the fan widens with lead time, so the same question has a very different answer at three days and at fifteen. The agent has to carry the "as of" timestamp through every step, and the crisis prompts in Chapter 10 all pin one.

The Shapes, Concretely

# what each environment tool hands back to the agent (schematic)
weather  = {"location": (lat, lon), "hourly": [{t, temp_c, precip_mm, wind_kph}] * 240}
flood    = {"gauge": id, "severity": "high", "probability": 0.71, "inundation": Polygon}
cyclone  = {"init": "2024-09-23T12:00Z", "members": [           # 50 of these
              {"track": [(lat, lon, t)], "wind_kt": [...], "r_hurricane_km": [...]}]}

# the question "counties in hurricane-force winds" against an ensemble
band = union(buffer(m.track, m.r_hurricane_km) for m in cyclone.members)   # one polygon
hit  = [c for c in counties if c.geometry.intersects(band) and c.pop > 20000]
Why the environment silo is the hardest to join by hand: its objects are polygons and time series that change every few hours, and its answers are probabilistic. A human analyst tends to collapse the ensemble to its mean track and lose the tails. An agent that can iterate over members, as the code above does, keeps them. That is a large part of why the agent beats plain Gemini by the widest margins on cyclone prompts in Chapter 10: 0.93 versus 0.23 on the typhoon scenario.
The cyclone model returns 50 complete scenarios rather than one best track with a confidence interval. What does that representation make possible that a single track with error bars does not?

Chapter 7: The Synergy Lab

This is the chapter the whole paper turns on. Each silo's model is good. The claim is that they are complementary: that a landscape embedding and a population embedding for the same census tract know different things, so a model fed both beats a model fed either. The evidence is two tables covering 41 target variables, and the recipe that produced them is simple enough to rebuild in an afternoon.

The Recipe

Take two public embeddings. AlphaEarth Foundations summarizes optical satellite imagery, radar, climate simulations and more into a 64-band embedding at 10-meter resolution, one image per year, available in Earth Engine. Population Dynamics Foundations gives one vector per US postal code. Neither is at census-tract resolution, and tracts are what FEMA and the CDC report on. So both have to be moved:

1 · AlphaEarth 10 m pixels → 640 numbers per tract
Take the 2023 image, 64 bands. For each band, collect every 10 m pixel inside the tract and summarize its values as a fixed 10-bin histogram. 64 bands × 10 bins = 640 features per tract. The histogram keeps the distribution (a tract that is half forest, half parking lot is not the same as one that is uniformly suburb), which a mean would erase.
↓
2 · PDFM postal codes → one vector per tract
Postal codes and tracts overlap irregularly. Use the 2020 land-area overlap from census.gov to compute, for each tract, an area-weighted average of the postal-code vectors that cover it.
↓
3 · Gradient-boosted decision trees, fixed hyperparameters
Concatenate the features. Split tracts by county: tracts in 80% of counties train, tracts in the other 20% test, so a model never sees the neighbor of a test tract. Train one GBDT per target variable, same settings for all 41. Report R² on the held-out counties.
From 64 Bands to 640 Features

One tract, one AlphaEarth band. The pixels inside the tract boundary each carry a value for this band; press Histogram to bin them into ten fixed buckets. That row of ten counts is what the boosted trees see for this band. Step through bands to fill the 64×10 feature block on the right.

Why a histogram and not a mean? Because FEMA risk is about extremes. Coastal-flooding risk depends on how much of a tract is at sea level, not on the tract's average elevation. A mean of the elevation-like bands would say "moderately low"; a histogram says "12% of pixels in the lowest bin," which is the number that matters. The feature construction is a modeling decision disguised as plumbing.

The Synergy Lab: 41 Targets, Three Feature Sets

Every number here is from the paper's Tables 9 and 10: R² on census tracts in the 20% of held-out counties. Pick a target, toggle feature sets, and watch which silo carries which variable. Then press Sweep all to see the whole picture at once: where fusion wins, by how much, and the two places it does not.

Reading the Sweep

On the twenty FEMA National Risk Index targets, PDFM alone averages R² 0.54, AlphaEarth alone averages 0.54, and both together average 0.60, an 11% relative lift. But the average hides the structure, and the structure is the lesson. Sort the targets by which silo wins alone:

FEMA targetPDFMAlphaEarthBothWho knows this?
Earthquake risk0.720.840.86The landscape: faults, terrain, building density from above
Social vulnerability score0.300.440.48Hard for both; imagery surprisingly ahead
Drought risk0.670.500.67Population: agriculture, search patterns; imagery adds nothing
Wildfire risk0.770.540.76Population wins; fusion slightly hurts
Coastal flooding risk0.330.240.38Both weak, fusion best: coast + who lives there
Overall risk score0.530.490.60The composite gains most from fusion

FEMA's own definition explains the pattern. A risk index multiplies expected annual loss (a physical quantity: how likely and how severe the hazard is, which imagery, terrain and coastline encode) by social vulnerability and divides by community resilience (human quantities: income, age, density, which the population embedding encodes). Risk is bimodal by construction, so a model that sees only one mode is capped. The overall Risk Score, the most composite target, gains the most from fusion: 0.53 and 0.49 alone, 0.60 together.

On the 21 CDC health indicators the story is cleaner: fusion wins on every variable. The means are 0.42 for PDFM alone, 0.56 for AlphaEarth alone, 0.60 for both.

What the paper does not say clearly, and you should notice: the prose reports the CDC gain as "7% over Population Dynamics Foundations alone and 43% over AlphaEarth alone." Table 10's columns read 0.42 (PDFM), 0.56 (AlphaEarth), 0.60 (both). From those column values the lift is 43% over PDFM and 7% over AlphaEarth, the reverse of the sentence. Either the prose swapped the names or the column headers did. The numbers themselves are unambiguous; the lesson is to trust the table, and to compute the ratio yourself rather than quoting a sentence. It also means the health result is not "the population model is best and imagery helps a bit." It is the imagery embedding carrying most of the health signal, with population adding 7% on top, which is a more surprising and more interesting claim.

Why Boosted Trees, and Why Fixed Hyperparameters

A gradient-boosted tree ensemble fits a sequence of small decision trees, each one correcting the residual error of the ones before. It is the standard tool for tabular data with hundreds of numeric features and a few thousand rows, which is exactly this shape (about 700 features, tens of thousands of tracts). It needs no feature scaling, handles the histogram bins and the embedding dimensions in the same way, and gives a direct measure of which features it split on.

The fixed hyperparameters are the point. The authors could have tuned per target and squeezed out a few more points of R². Instead they held everything constant across 41 targets so that the only variable is the feature set. The result "suggests that a simple modeling setup can benefit from the gains of combining embeddings without tuning to each individual label," which is the property an agent needs: in Chapter 9, the Spatiotemporal Model Training expert trains models like this on the fly for whatever variable a user names.

# the whole synergy experiment, in the shape the paper describes
ae   = alphaearth_histograms(tracts, year=2023, bins=10)      # [T, 640]  64 bands x 10 bins
pd   = area_weighted_pdfm(tracts, overlap_2020)               # [T, D]    zip vectors -> tract
X    = {"pdfm": pd, "alphaearth": ae, "both": np.hstack([ae, pd])}
train_ct, test_ct = split_counties(counties, test_frac=0.2)     # split by COUNTY, not by tract
for target in fema_20 + cdc_21:
    for name, feats in X.items():
        m = GradientBoostingRegressor(**FIXED)              # same settings for all 41 x 3 fits
        m.fit(feats[in_(train_ct)], y[target][in_(train_ct)])
        r2[target][name] = m.score(feats[in_(test_ct)], y[target][in_(test_ct)])
The county split is the quiet safeguard. Adjacent tracts are nearly identical, so a random tract-level split would let the model memorize a neighborhood and call it generalization. Holding out whole counties forces every test tract to be somewhere the model has never been. Every R² in this chapter is an out-of-county number.
On wildfire risk, PDFM alone scores R² 0.77, AlphaEarth alone 0.54, and both together 0.76. What is the right reading of that row?

Chapter 8: Forecasting With Covariates

Chapter 7 fused two static embeddings to predict a static number. The environment silo adds time, and time changes the question from "what is this place like" to "what happens here next week." Two case studies show the same move with a clock attached: take a forecast, and give it what the other silos know.

Case 1: Hurricane Ian, Three Days Out

Bellwether, a Google X project, built a model that answers one question per building: will this building be wind-damaged by this storm? The training data is 1.5 million data points across tens of thousands of residential and commercial properties, six years (2017–2022), 22 hurricanes, with damage labels and property details (material, roof type), plus storm information from the HURDAT2 archive, building heights from Open Buildings, and the Population Dynamics embedding of the building's region. Six of the most recent storms are held out. The model is, again, gradient-boosted trees.

Here is the trap in that dataset, and the paper walks straight into it on purpose: fewer than 4% of properties are damaged in any storm. A model that says "undamaged" to everyone scores 96% accuracy. So the paper reports the numbers that matter, the ones for the minority class:

Held-out test, observed tracksRecallPrecisionF1
Class 0: undamaged buildings0.980.980.98
Class 1: damaged buildings0.600.570.59

Recall 0.60 means the model finds six of every ten buildings that will actually be damaged; precision 0.57 means when it flags a building, it is right a little more than half the time. Modest, honestly reported, and useful, because the operational question is not "which building" but "how many," and counts average out per-building errors.

Then comes the forecast. Replace the observed track from HURDAT2 with the cyclone model's forecast track, issued September 26th, 2022 at 00 UTC, three days before Ian's landfall. There are fifty forecast members, so run the damage model fifty times and collect fifty counts.

Fifty Storms Through One Damage Model

Each ensemble member is pushed through the per-building damage classifier, producing one predicted count of damaged buildings. Press Run ensemble to build the distribution member by member. The dashed line is the ground truth for this dataset, 2,496 buildings; the mean of the fifty forecasts landed at 2,575. Below, the fifty precision/recall/F1 values for the damaged class spread into box plots.

2,575 predicted against 2,496 observed is a 3% error, three days before landfall, and the spread of the distribution tells a relief planner how much to hedge. The paper is careful here: this is one storm at one lead time, the dataset is a subset of buildings in the affected zone (the true damage count was far higher), and some "wind damage" labels are probably flood damage. Read it as a demonstration of the shape of the pipeline, not as a validated forecasting product. The shape is the point: an ensemble forecast, a per-unit model fed static embeddings, and a distribution out.

Case 2: Cholera in the Congo, With Weather and Population as Covariates

The Democratic Republic of Congo has endemic cholera, and with WHO's African regional office the authors built a sub-national forecaster for weekly case counts in 24 provinces. The base model is TimesFM 2.0, a decoder-only foundation model for time series that can forecast a new series zero-shot, with no training on it. The experiment adds covariates in layers:

A Forecast, Then the Other Silos

A synthetic weekly case series for one province with a rainy-season swell. Toggle the layers: TimesFM alone extrapolates the history; adding the weather forecast (rain and temperature drive transmission) lets it anticipate the swell; adding the province's PDFM embedding shifts the level to match what the province is like. The error readouts follow the paper's reported improvements.

LayerWhat it addsReported effect
Prophet / SARIMAXClassical statistical baselinesThe reference
TimesFM 2.0 zero-shotFoundation-model extrapolation of the weekly counts alone35.4% lower RMSE than Prophet at a 20-week national horizon
+ dynamic weather covariatesForecast precipitation and temperature, the known drivers of transmissionA further 32.7% improvement
+ PDFM static covariatesOne embedding per province, describing what it is like5% lower RMSE at the regional level versus TimesFM alone

The ordering of the gains is instructive. The biggest single step is swapping a classical model for a time-series foundation model, before any Earth AI signal is involved. The second biggest is the environment silo: weather is causal for cholera, so a forecast of rain is a forecast of cases. The population embedding adds the least, 5%, and only regionally, which makes sense: a static description of a province helps set its level but cannot tell you about next month's outbreak. Each silo contributes what it knows, and the paper reports the small gain as a small gain.

The recipe, generalized: a forecast model plus static covariates from the population silo plus dynamic covariates from the environment silo. It is the same pattern the dengue validation in Chapter 5 found (PDFM helped TimesFM most at 6–12 month horizons), and it is what the agent's Spatiotemporal Model Training expert automates: "forecast unemployment for October in every county" becomes TimesFM on the county series, with the county embedding attached.
For Hurricane Ian the paper reports a mean predicted count of 2,575 damaged buildings against 2,496 observed, but also that the per-building damaged-class F1 is only 0.59. How can both be true?

Chapter 9: The Agent

Everything so far is a specialist. Chapter 9 is the generalist that hires them. The Geospatial Reasoning Agent is built on Google's Agent Development Kit (ADK) with Gemini 2.5 Flash and Pro, and its design is a registry of capabilities organized by domain, plus a loop that plans, acts and checks. It is worth being precise about what is a tool and what is an expert, because the paper found the difference matters.

Experts and Toolsets

Some domains "succeeded as a simple list of tools." Others "needed more steering and guidance," so they were wrapped as expert sub-agents: a sub-agent has its own prompt, its own judgement about which of several data sources to use, and returns a synthesized answer rather than a raw payload.

CapabilityKindWhat it does
DemographicsExpertChooses and retrieves the right statistical variables from Data Commons for a natural-language question
WeatherExpertFetches and interprets historical (ERA5 via Earth Engine) and forecast weather, cyclones and floods
Spatiotemporal Model TrainingExpertTrains lightweight models on the fly on PDFM and TimesFM for interpolation, super-resolution and forecasting of a user-named variable
Earth EngineExpertFinds the right public dataset and band in the Earth Engine catalog by searching it, then fetches
SearchExpertA fallback: Vertex AI grounding with Google Search when nothing specialized has the fact
Location & PlacesToolsetTurns names (a county, "hospitals in Louisiana") into polygon geometries via the Places and Places Aggregate APIs; the foundation for every spatial intersection
Remote Sensing & ImageryToolsetThe Chapter 2–3 models: open-vocabulary classification, retrieval and detection, with cached image embeddings so repeat queries over the same area are cheap
GeneralToolsCode generation for custom analysis, Earth Engine access, Google Search, Google Cloud services

The paper's stated reason for the split is worth quoting in paraphrase: Gemini's reasoning is already good when all the relevant data is handed to it in natural language. Where it needs help is access to specialized data, special handling of non-language data such as geometry, and orchestration when many tools have to be sequenced and their outputs joined. Experts exist for the domains where the second problem is worst.

The Loop

The operational framework is a closed cycle with three stations. Think & Plan: decompose the query into sub-tasks and pick a capability for each. Act: data operations, model inference, or model training, whichever the sub-task needs. Reflect & Recover: check the result; if a tool failed or returned nothing usable, revise the plan and go around again. The cycle runs until a final, grounded response exists, and the response carries its provenance: which data source each number came from.

Hurricane Helene, Step by Step

The paper's worked example, as the agent runs it. The prompt: "As of September 23, 2024 12:00 UTC, what did Google's cyclone model say the predicted path of Hurricane Helene was? Get the list of counties in Florida with population > 20,000, and filter the list to the ones predicted to experience hurricane-force winds." Step through the plan; the map on the left updates with each tool's output and the loop on the right shows which station is active.

Trace the data types through those steps, because they are the whole design in miniature. Step 1 returns an ensemble: fifty tracks with wind radii. Step 2 collapses it to one polygon, the union of hurricane-force wind bands. Step 3 returns a table from the Census via Data Commons: county name, population. Step 4 returns polygons from the Places toolset, one per county. Step 5 is pure geometry: intersect polygons with the wind band, filter by population, and four counties survive. Step 6 puts them on the map in green and writes the sentence. No single model did anything hard. The agent's contribution was knowing that a cyclone forecast, a census table and a boundary service are three calls that compose.

Where the base model fails is exactly the seam. Plain Gemini 2.5 Pro with search and maps grounding, asked the same thing, has to read about Helene's forecast cone in news coverage and guess which counties it covered. It cannot fetch the members, cannot compute the union, cannot intersect. In the Chapter 10 benchmark, the hurricane-plus-demographics prompt scores 0.45 for base Gemini and 0.80 for the agent, and the typhoon-plus-location prompt scores 0.23 versus 0.93. The gap is geometry, not intelligence.

The Provenance Habit

Look at the agent's answer in the paper's Matanuska-Susitna flood example (Appendix A.3.3): a table of three zip codes with income and share over 60, then a numbered reasoning trace (flood forecast fetched; zip codes listed; polygons fetched; spatial intersection performed; statistics retrieved per zip), then a Data Provenance block with a source per column. That structure is not decoration. It is what the rubric-based grader in Chapter 10 checks for, and it is what makes an analyst trust the list enough to act on it. Base Gemini's answer to the same prompt names no source for its flood data, admits some percentages "were estimated," and lists eight zip codes where the authoritative flood polygon touched three.

# the Helene query as the agent's plan (schematic of ADK tool calls)
fc   = weather_expert.cyclone_forecast(storm="Helene", as_of="2024-09-23T12:00Z")   # 50 members
band = union(buffer(m.track, m.r_hurricane_km) for m in fc.members)        # Polygon
pop  = demographics_expert.query("population of each county in Florida")          # table via Data Commons
geom = places.polygons([c.name for c in pop if c.population > 20000])              # {name: Polygon}
hit  = [n for n, g in geom.items() if g.intersects(band)]                             # 4 counties
answer(table=hit, map=[band, *geom.values()], provenance={"forecast": fc.source, "population": pop.source})
Why does the paper implement some domains as "expert" sub-agents with their own prompts rather than as plain tools?

Chapter 10: Scoring an Agent

How do you grade an agent that returns tables, maps and paragraphs? The paper builds two benchmarks with two different grading philosophies, and the second one is the more interesting because its limits are stated so plainly.

Benchmark 1: 100 Fill-in-the-Blank Questions

A hundred questions across four domains (Places 30, People & Communities 40, Weather & Environment 20, Multiple Domains 10), half descriptive-and-retrieval and half analytical-and-relational, every one answerable from public data and every one with a manually verified gold answer. Every question is phrased as fill-in-the-blanks so the grader can ignore prose: "Respond by filling in the blanks: There are __ botanical gardens in Lincoln, Nebraska." A subject-matter expert reviewed each question for being unambiguous, time-invariant in the short term, and useful to a real analyst.

Scoring: text blanks get ROUGE-L F1 against the gold. Numeric blanks get a clamped absolute percentage error:

APE = |predicted − gold| / |gold|
score = 1.0 if APE ≤ 0.10;   0.0 if APE > 1.00;   1 − APE otherwise

The 10% dead zone is deliberate: businesses open and close, statistics get revised, two sources disagree at the margin, and none of that should cost points. The clamp at 100% error is a floor: being off by a factor of two is as wrong as being off by a factor of ten. A question's score is the mean over its blanks; the benchmark score is the mean over questions, averaged across five runs with a 95% confidence interval.

The Scoring Lab

Left: the numeric scorer. Drag the prediction against a fixed gold value and read the score off the curve; note the flat top and the hard floor. Right: the paper's Table 11, drawn as bars. Toggle the category to see where the gap between the agent and plain Gemini opens up, and by how much.

Q&A benchmark, mean of 5 runsPlaces (30)People & Communities (40)Weather & Env. (20)Multiple (10)Overall (100)
Gemini 2.5 Flash + built-in tools0.350.460.350.280.39 ± 0.02
Gemini 2.5 Pro + built-in tools0.450.540.540.380.50 ± 0.01
Geospatial Reasoning Agent0.730.910.710.980.82 ± 0.02

The baselines are not straw men. Both Gemini agents are also built in ADK, with Google Search grounding, Google Maps grounding and code execution enabled, and are prompted to always give a best estimate and cite sources. They can look things up and run code; what they cannot do is call the models. Split by category, the agent beats Pro by 37% on descriptive questions (0.91 versus 0.67) and by 124% on analytical ones (0.74 versus 0.33). The harder the join, the wider the gap, which is the same pattern the Chapter 0 widget predicted.

Benchmark 2: Ten Crisis Scenarios, Graded by Rubric

The crisis prompts (Chapter 9's Helene query is a cousin of prompt 4) cannot have a gold answer, because they involve forecasts that change with every run and free-form outputs. So each prompt gets a rubric: a list of criteria such as "authoritative flood data was used," "there was a process to precisely identify which zip codes to include," "the response is an answer, not a plan for how someone could answer." An autorater built on Gemini 2.5 Flash reads the response against the rubric, writes a numbered justification per criterion, and issues a five-point verdict (not at all, somewhat, moderately, mostly, completely) mapped to 0, 0.25, 0.5, 0.75, 1.0. Each prompt runs ten times for each system.

#CrisisCapabilities testedGemini 2.5 ProAgent
1FloodFlood forecast + Places0.180.80
2FloodFlood forecast + Demographics0.280.80
3TyphoonCyclone forecast + Location0.230.93
4HurricaneCyclone forecast + Demographics0.450.80
5DroughtWeather forecast + Earth Engine0.580.88
6Extreme heatWeather forecast + Demographics0.761.00
7FireRemote Sensing + Places0.400.84
8FloodRemote Sensing + Demographics0.350.88
9HurricaneTimesFM + Demographics0.00*1.00
10HurricanePDFM + Demographics0.250.91
Mean (excluding 9)0.38 ± 0.170.87 ± 0.14

Row 9 is starred because base Gemini "was consistently unable to find authoritative labor statistics for 2024," so it could not attempt the question; the authors exclude it from the mean rather than count a free zero. Notice also that the base model is not hopeless: on extreme heat (weather forecast plus demographics, both of which search can partly reach) it scores 0.76. The gap is widest exactly where a polygon or a model call is unavoidable.

What the Rubric Cannot See

The paper devotes a full subsection to the limits of its own grading, and it is the most reviewer-like passage in the document. Two problems.

Consistent reasoning, inconsistent outputs. On the Maui fire prompt the agent's process was stable across ten runs, but the list of damaged schools and stores changed whenever the remote-sensing model flagged a different set of burned tiles. A rubric grades the process, so it cannot catch that the answer is unstable; only a gold answer could, and there is none.

Verifying methodology. For a widely covered event, a response might say it analyzed satellite tiles while actually synthesizing news reports found by search. The rubric checks for a claim of authoritative data use; it cannot easily check that the claim is true. The authors found this "particularly evident with high-profile events," where prompt refinement revealed that an apparently first-hand analysis was secondary sourcing.

Why this matters beyond this paper: every agent benchmark that grades with an LLM judge and a rubric inherits both problems. The Earth AI authors' proposed fixes are the honest ones: gold answers wherever a deterministic one exists, and human expert review alongside the autorater everywhere else. Their stated plan for future work is a broader task set, better rubrics, and human review, plus tests of out-of-distribution queries and of what happens "in scenarios where specialized tool orchestration fails," which the current benchmark does not probe at all.
A numeric blank has gold value 100. Under the paper's scorer, which of these predictions all receive exactly the same score?

Chapter 11: Limits & Beyond

A paper this broad earns trust by saying where it stops. Each of the four sections has a stated limitation, and each one is a research direction in disguise.

ComponentStated limitationWhat it implies
Remote Sensing FoundationsHigh-resolution RGB only; no temporal downstream tasks evaluated; oblique, multispectral and hyperspectral imagery not yet supportedChange detection over time (the thing crisis response most wants) is not benchmarked; a Sentinel-2 multispectral user still needs other backbones
Population Dynamics FoundationsTemporal look-back is a few months; more signals and finer granularity wanted; interpolation tests hide regions at random, not in the systematic gaps real outages createThe 0.85 R² is an optimistic bound for the case where a whole region's data goes dark at once
Model synergyAggregating 10 m AlphaEarth embeddings to tracts by histogram is crude; the fused model needs training per target variable; no single unified "meta-Earth" model exists yetThe obvious next paper: one model trained jointly on imagery, environment and population signals with a shared representation, replacing late fusion
Reasoning agentBenchmarks test the agent's designed capabilities, not out-of-distribution queries or tool-orchestration failures; rubric grading has the two blind spots of Chapter 10The 0.82 and 0.87 are in-distribution numbers; the hard cases are exactly the ones not measured
The unstated assumption to keep in view: nearly every data source in this system is Google's (Maps, Search Trends, busyness, Earth Engine, the weather and flood and cyclone models, Data Commons). The architecture is general, and the embeddings and several APIs are public, but reproducing the Population Dynamics embedding from scratch is not possible outside the company, because the search and busyness signals are not. When you read "foundation model" here, read "foundation model over a proprietary sensor."

The Cheat Sheet

ThingDefinition in one lineThe number to remember
RS-SigLIP2 / RS-MaMMUT400M-param contrastive VLMs fine-tuned on RS-Landmarks (18M Gemini captions), RS-WebLI (3M mined), RS-Global (30M native-res)RESISC45 zero-shot 72.40 → 80.13; RSITMD retrieval I2T 24.62 → 43.14
RS-OWL-ViT-v2 + FLAMEOpen-vocabulary detector, 2-stage cooldown fine-tune; FLAME picks uncertain-then-diverse examples for few-shotDOTA mAP 13.77 → 31.83 zero-shot → 53.96 with 30 labels
RS-Global MTP backboneMAE on 300M images, then multi-task (cls/seg/det) with alternating gradient descent+14.93% avg over ImageNet ViT-L; FMoW 81.70, FLAIR 65.72, DIOR AP50 85.50
Population Dynamics FoundationsGNN over regions; Maps + Knowledge-Graph search entities + busyness + weather → one vector per postal code; 17 countries; monthlyR² 0.85 across 17 countries; France from 7 neighbors: 0.52 GDP vs 0.00 for IDW
Environment modelsWeather (hourly to 240 h), floods (severity + probability + polygon), cyclones (50 members, 15 days)50 scenarios, not one track
Synergy recipeAlphaEarth 64 bands × 10-bin histograms (640) + area-weighted PDFM → GBDT, county-level splitFEMA R² 0.54 / 0.54 / 0.60 (+11%); CDC 0.42 / 0.56 / 0.60
Forecast + covariatesTimesFM 2.0 zero-shot, then weather (dynamic) and PDFM (static) covariatesCholera: −35.4% RMSE vs Prophet, a further −32.7% with weather; Ian: 2,575 predicted vs 2,496 damaged
Geospatial Reasoning AgentADK + Gemini 2.5; five experts, two toolsets; Think & Plan → Act → Reflect & RecoverQ&A 0.82 vs 0.50; crisis rubric 0.87 vs 0.38
ScoringROUGE-L for text; 1.0 within ±10%, 0 beyond 100%, else 1 − APE; rubric → Likert → {0, .25, .5, .75, 1}Five runs, 95% CI reported

Connections on This Site

The pieces of Earth AI are each taught from zero elsewhere here. The contrastive image–text machinery of Chapter 2 is Contrastive Learning & CLIP, and the sigmoid-loss variant this paper fine-tunes has its own Veanor in EVA-CLIP's neighborhood. Masked autoencoding, the Stage 1 backbone objective, is worked through in Self-Supervised Learning and, for the audio cousin, in the Audio-MAE Veanor. The graph neural network that produces population embeddings is Graph Neural Networks plus the CS224W sequence starting at GNN I. FLAME's second move is k-means. Detection and segmentation heads are in Detection & Segmentation. The agent loop, tool use, and how to evaluate an agent with rubrics are covered by Agents & Tool Use, Agent Evaluation and the agent-evaluation survey Veanor. Knowledge distillation, the reason a small backbone can inherit a big model's competence, is Knowledge Distillation.

What to Try Yourself

The synergy experiment is reproducible with public pieces. AlphaEarth embeddings are in Earth Engine; the US Population Dynamics embeddings and code recipes are on GitHub at google-research/population-dynamics; FEMA's National Risk Index and the CDC PLACES tract-level indicators are open downloads. Build the 640-feature histogram block, area-weight the zip vectors to tracts, split by county, and fit one boosted-tree model per target. If you get an 11% lift, you have reproduced the central claim of a fifty-author paper on a laptop. If you do not, you have found something worth writing up.

The authors name a "single, unified meta-Earth model" as the significant future advance. What specific technical hurdle do they say stands in its way?