Also published on LinkedIn.

Those of us at the intersection of EO and AI still have a lot to bring, but further maturing a separate domain might not be the most impact-driven path.

This is the article form of the talk I gave at #ECCV conference workship on Conservation and Climate change. Thank you Climate AI Nordics for the invitation. Hope I was able to sir the pot enough.

Is geodata just data?

We geo experts adopt most frontier techniques fast — SAM (“Segment Anything Model”) was applied to satellite imagery within three weeks, with fantastic applicability. Yet there is no evidence that the general-purpose frontier models (Gemini, GPT, Claude, DeepSeek, Kimi, GLM) treat Earth observation in any of the specific ways we in GeoAI do: feature engineering around metadata, rasters, projections, bands, and sensor physics. Models that do accept images expose only an RGB image interface, and while they can appear well versed and able to generate GIS code, the primitives that define GeoAI (e.g. geoattention) are absent from their published core architectures. I have reviewed eighteen frontier-lab model cards and technical reports from this year. Not one documents an Earth-observation input or an Earth-observation benchmark. This is absence of evidence, not evidence of absence, but it does show that geo is not a publicly stated architectural priority.

Even the newest wave of AI focused on “world models” might look like a turn toward Earth data. But read the actual release pages across twenty world models: their inputs are video, text, actions, and lately 3D scenes; their stated purpose is intuitive physics, ego views, simulation, robotics, and autonomous driving. Again, not one lists satellite or Earth-observation data, or even Earth context, as a consideration. They simulate the world with real data, without anchoring it to the Earth itself.

Talk slide surveying frontier AI and world models, with none listing Earth observation as an input.

Google is the exception among those labs. While Gemini still lacks geo priorities, Google has a dedicated GeoAI programe, “Earth AI”, which houses the AlphaEarth model, one of the leading models. In the last two weeks, Google also released the Planetary Prediction Engine and WeatherNext 3; Implicitly or explicitly, Google seems to follow the very thesis I am arguing in this talk, which is that the best path for geo and Earth in AI is to 1) maximize the use of general AI models in geo applications, 2) constrain geo models as tightly as possible and leave the reasoning to general models, and 3) figure out how to anchor general AI on Earth when it makes sense. I return to this later.

What stays out when models do not treat EO as distinct?

My coworker Konstantin Klemmer, with Rolf, Robinson and Kerner, argues that satellite data is a distinct modality in machine learning. I agree completely with the measurement claim: spectral bands, acquisition geometry, scale, time, and spatial autocorrelation are not ordinary natural-image context. My take is that a distinct modality does not justify keeping a separate model family, nor does it represent an isolated kind of intelligence.

This is not a claim that all images are the same. Specialist models win when input and task match, the more narrowly they specialize the more clearly they win, and they will keep winning there. The job for GeoAI is to understand what is distinct, and to make those signals usable and useful to general reasoning systems.

Consider the epistemological difference. Words in language are symbols; their meaning emerges from relational context. These sequences of symbols capture all of history (by definition): law, property, economics, facts, fiction and intention. Today’s frontier transformers learn by attending across that vast web of human meaning. When an LLM encounters e.g. “Venice”, it draws on its history, engineering, tourism, and sea-level rise. LLMs encapsulate history itself.

Satellite view of the Haiti–Dominican Republic border, showing contrasting vegetation cover.

the text “Haiti — Dominican Republic border” should converge to the above image, but they mostly complement each other

Earth observation units are fundamentally different. They are explicitly records: photons, emitted radiance, backscatter, and elevation at a specific coordinate and time. Unlike words, their meaning maps directly to physics. Those measurements, alongside the metadata of when and where they were taken, are the unit of data for GeoAI models. EO data is thus extremely good at encoding the physical context of what is happening, where and when. This is why imo GeoAI models, often orders of magnitude smaller and trained with far less compute and data, still perform much better on some matched EO tasks.

But EO sensors are blind to many other drivers that text can encode: legal tenure, property rights, environmental regulation, and political incentives. A sensor may detect a clearing in a forest, but it cannot read the timber concession permit or the land title. It therefore seems clear that general AI can access a much wider stated context than EO alone. Further, anything EO sees could be described in text and brought into context. An image might indeed be worth a thousand words, but a thousand words in context carry far more of the drivers than an image can.

That is the central thesis of this talk: there is no separate EO intelligence, just a specific type of data that needs specific means to be used well. AI without EO could plausibly reach where EO gets, while AI models built only on EO will have a much harder time reaching the drivers of context than AI models without it.

I argued for this conceptual approach in a piece for Davos this January: use AI to make Earth patterns searchable and shared baselines usable without specialist teams. This talk makes its architectural consequence explicit: Earth observation is a distinct measurement, but not a separate intelligence.

Why GeoAI became separate

If that is true, why do we have a separate field at all? We have built a whole GeoAI field, with our own methods, benchmarks and models. Our field performs clearly differentiated and specialized GIS and EO tasks with heft, maturity and quality. It works. But it is also extremely niche. The whole GIS field has struggled for decades to widen adoption and lower skill and budget barriers, and in general longs for much more adoption and impact.

I have argued in the past that this eternally stunted reach in geo stems from the origins, and even today the major drivers of revenue, of the geospatial field: defense and the extractive industries, with large budgets, highly skilled people, and a focus on narrow, high-value problems. In these fields it continues to perform very well. This keeps pushing the tip of the spear, while creating no incentive to lower the barriers to entry for the long tail of impact opportunities, including climate and conservation. Geopolitically, we are also seeing increased military interventions, sovereign objectives and a “dual use” posture on one side, and divestment from voluntary carbon markets and ESG strategies on the other. This reinforces the very drivers that disincentivize the democratization of geo and GeoAI. This is not a critique or judgment. Geo work extremely well for the ones that depend and deploy it the most.

In any case, the gap between AI and geoAI is so wide that we do not even use the same benchmarks. GeoAI largely does not use AI benchmarks, or benchmark frameworks, but creates its own, on the basis of differentiated needs. Even worse, Corley, Kerner, and colleagues (arXiv:2605.12678) recently audited 152 geospatial foundation model papers and cataloged 401 distinct benchmarks. In 46 direct comparisons of the same nominal model and dataset, reported scores disagreed by more than ten percentage points; 39% of papers released no public weights. When every team creates its own isolated evaluation, we are not measuring the signal of progress toward planetary outcomes; we are chasing the signal of novelty.

For me, the question is not how to improve or protect the fork of AI we call GeoAI. It is how to merge GeoAI back in, contribute its strengths upstream, and use every kind of AI not to advance benchmarks, but to advance climate and conservation goals.

Where impact actually breaks

Let’s step back and look at the entire theory of change. Is the primary bottleneck in climate and conservation really the third decimal point of precision, or architectural tweaks to a seasonally invariant attention?

Scientific diagnosis has been precise, but largely ineffective.

In The Environmental Republic, Giulio Boccaletti , former Chief Strategy Officer at The Nature Conservancy, wrote: “Scientific diagnosis has been precise, but largely ineffective.” That is not an attack on science; it is a warning against mistaking a diagnostic tool for an outcome. As he says, the authority of science comes from who is asking the questions, and its ultimate utility from what happens with its output. It is what I argued seven years ago in my book Impact Science: better evidence does not automatically produce better decisions.

Consider the study by him, and colleagues in Conservation Biology (2024), evaluating 3,179 land parcels protected by The Nature Conservancy across three US regions over nearly three decades. I am sure very complex models drove those decisions, and that today we have even better ones. Alas, the majority of the portfolio showed no evidence of avoided conversion on average compared to matched parcels. Why? Because the easements were secured on land that was not under real threat of development. The failure was in the institutional capacity to deploy capital where it would have had the most impact, rather than where it could.

Environmental impact depends on more than data and models: someone must integrate the output into a real decision, maintain the system and check whether it worked. A better foundation model would not have changed where TNC bought easements. We have more data than ever before, we have better models than ever before, but we lack similar progress in our capacity to drive impact. In fact, the abundance of data and outputs seems to drive more paralysis. As James Gleick puts it in his book The Information, “when information is cheap, attention becomes expensive”. The current challenge is less the data and the models, and more getting and shaping the attention of the stakeholders who could make any difference than on any model.

So how does GeoAI fit? It may seem odd to ask a technical field to reach both upstream and downstream of science, but that is precisely the point. AI can help make better models — even better frontier models. I will return to what I think it’s the exact mechanism to anchors AI reasoning on Earth (embeddings injection!) , but the bigger opportunity is helping people understand, integrate and act on science. Empower stakeholders; do not just throw decimal points at them from a paper they cannot open or a model they cannot run.

One case: Nepal

Flood debris and a river channel beside a settlement in a steep Himalayan valley in Nepal.

Even if you reject everything I have said about GeoAI and general AI, the more important problem remains: neither produces impact by itself. The two halves of this talk (geoAI and impact) are one problem. The niche adoption that has held geo back for decades is the same wall that stops evidence from reaching a decision, and general AI is the first technology that can help cross to both sides of that wall.

Two weeks before this talk, on 26 August, ice and rock came off a glacier in Langtang and dammed the Lhende Khola; then the dam let go. No models helped prevent, alert or reduce the risk. Just days ago, a geophysicist at Virginia Tech went back to the free Sentinel-1 radar archive. The slope had been creeping downhill at about ten millimetres a month since January, and the creep was accelerating in the days before it let go. Optical imagery showed nothing; the radar had seen it… but no one was looking. One could say that monitoring everywhere every time is nearly impossible… and I’d counter nobody needs to look everywhere, local teams everywhere monitor their local needs. That is indeed possible more and more as we develop better data and tools. And for Nepal, they had been there before.

On 16 August 2024, a rock avalanche fell into a glacial lake above Thame, a hundred kilometres east in the Khumbu; the wave breached a second lake and the flood hit the village. Neither lake was on Nepal’s list of dangerous lakes. The Government of Nepal–ICIMOD assessment came afterwards, from Planet imagery reached through the NASA–USAID SERVIR programme. The gap was which lakes counted and who was tasked to look. The data, again, had already been captured and sat in satellite archives. And then again on Imja, a glacial lake a valley over. Nepal’s hydrology department, with UNDP, lowered the lake 3.4 metres and put sirens in six villages along fifty kilometres of valley. Then the project closed in 2017. This April, the department that owns the sirens told the BBC it could not say whether they still work. The new programme is about fifty million dollars over seven years. It lowers four other lakes… none of them in the Lhende–Bhote Koshi valley that flooded two weeks ago.

Nepal does not tell us whether GeoAI or general AI is the better architecture. And I hesitated a lot to bring it on this talk because so much is still unknown, and so many needs to be done to help that is to distract the issue. But something it tells me, extremely loud, is that scientists cannot stop at the model. This is the Environmental Republic thesis again. GeoAI can improve the measurement, provenance, uncertainty and update cycle. General AI may help connect that evidence to forecasts, local language, plans and workflows. But institutions and communities must authorize warnings, maintain sirens, fund engineering, rehearse evacuation and learn from the outcome. The role of the GeoAI community is not to sell Nepal a heavier checkpoint. It is to empower those affected or benefited by it to make the Earth measurement reliable, legible and composable inside a whole system that is locally owned, so they can make their own decisions.

How we merge back

The transition is already underway. At LGND, we are building infrastructure that bridges both worlds: producing high-fidelity Earth representations, but coupling them more and deeper with frontier general-purpose reasoning models that operate under real-world constraints. We are not alone in recognizing this shift; Google’s design choices with the Planetary Prediction Engine reflect the exact same convergence. Rather than attempting to expand AlphaEarth into a monolithic, do-it-all geoAI system, they built an agent around a general-purpose language model, turning specialist representations like AlphaEarth and PDFM into modular tools on the agent’s menu.

What improves when models consider EO as Earth anchors?

General AI models are trained with vastly more data and compute than anything we can dream of in GeoAI. We will not out-compute the frontier. Not only that, Rich Sutton’s The Bitter Lesson teaches us that “general methods that leverage computation are ultimately the most effective.” In other words as much as our engineered geo ways do work today, AI throwing compute at it will eventually get there. A paper out of Berkeley and Stanford just last week is the latest proof: Hand-designed data agents beat a plain coding agent on “old” last year’s models; on this year’s, the general agent wins handcrafted experts on accuracy, time to completion and cost. And when latest models fail also teaches us: they failed when the context had changed since they were trained. Even the authors’ view was to build a context once and reused across every query. That paper wasn’t on Earth data but the application is exactly what we’ve proposed: Make Earth itself, through embeddings the complete, updated, comprehensive context. Scale can eat the geo in geoAI; but it cannot see Earth itself.

And there lies the bridge. Embeddings. Text can be anything; but Earth is one. Large, but one. Whenever one of these powerful models reasons about somewhere and some time, it seems sensible to anchor it in the available measurements of reality. And models reason in embeddings, so geoembeddings seem the best-digested artifact to serve as the bridge. This is probably the single most important thing frontier AI should understand about EO. Geoembeddings are dense artifacts that capture “all” actual semantics at once of a spatiotemporal location, and they can be aligned to the context space of AI models.

Embeddings are the missing context to ground AI on Earth

The exact mechanism is something we are actively investigating, but conceptually it stands to reason. If I speak about, say, Malmö today, the text embedding of Malmö might have learnt its context, culture and even geography, but bringing the real EO geoembedding of Malmö from, say, Sentinel-2 into available context will add, at once, what a many words cannot. It might see that it has been very dry, or wet, or the construction close to us. And like polysemy on a text token, it will only activate the relevant meaning in the current context. If information is cheap and attention expensive, an embedding is the cheapest way to help a reasoner’s attention.

DFR-Gemma makes this less hypothetical. They projected 330-dimensional Population Dynamics Foundation Model embeddings directly into a frozen Gemma 3 4B embeddigns space as continuous soft tokens. On complex multi-region tasks, these reached 0.72 accuracy, versus 0.46 for textual descriptions and 0.39 for a LightGBM baseline. They did not do satellite EO. Those embeddings mostly describe socioeconomic activity, amenities and weather, and the technical mechanism is designed to be as simple as possible. But it shows that structured geographic context can enter a general-purpose model through this interface.

Specialist GeoAI models will still win when the input and task is narrowly scoped to signals readably available on EO data (meaning not implicitly sensitive to the context LLMs bring).

Google’s Planetary Prediction Engine is aligned with the thesis of this talk. PPE is an LLM agent tasked with orchestrating geospatial data discovery, code execution, and model training. It is not a geo model; in fact, it is only given AEF and PDFM embeddings as a (forced) option among its tools. And PPE seems to work: across 21 public-health indicators, PPE improved mean R² from 0.600 to 0.768. Ironically, their ablations show that the embeddings added little: PDFM alone gives 0.597, adding AlphaEarth 0.618. Almost all the gain came from Census covariates the agent found on the web, and adding high-spatial resolution annual geoembeddings to zonal monthly statistics is doubly not the best choice. I would bet that simple linear probes on spatially and temporally aligned embeddings would have flipped the ablation. PPE also does not report compute or time spent per task, where embeddings would probably have performed much better than computing NDVI over the time series and regions.

So general AI is fantastically useful for thinking about geo, and embeddings can potentially fast-track and anchor that thinking. This is the convergence I mean. Not a separate GeoAI model trying to master general reasoning, but general AI orchestrating specialist tools and grounding itself in measurements of the Earth.

What I want people to leave with

  1. My hunch is that GeoAI will ablate away, and what remains will be Earth grounding for AI. General AI still has a lot of value left to squeeze, and we can expect better and cheaper general reasoning that will, almost accidentally, improve geo capabilities.
  2. The intersection of your skills is extremely rare and incredibly valuable.

Before you dive into a rabbit hole to gain a decimal point with a clever geo architecture, new data or a new benchmark, it might be worth stepping back. Take a moment to see how what already exists lands on the ground. If it doesn’t, ask why. Question assumptions, seek evidence, and understand the limitations. I will bet that in most cases the rest of your skills and experience place you, almost uniquely, to solve those real challenges. You come from there, or speak the same language, or know local groups, or have access to knowledge not publicly available, or share the same culture, …. Chances are the biggest impact comes from bridging the gap between the paper and the ground you stand on.