
Earth Science AI Is Moving From Perception to Scientific Workflows
A new arXiv preprint review maps how Earth science foundation models are evolving from task-specific perception toward multimodal reasoning and agentic scientific discovery—and catalogs more than 200 datasets and benchmarks along the way.
May 9, 2026
Imagine an Earth science system that does not stop at a single prediction. It retrieves related observations, calls scientific tools, and explains what it found. That is not a current operational system. It is the direction of travel described in an arXiv preprint review of Earth science foundation models.
Earth science is now a study of coupled spheres: atmosphere, hydrosphere, lithosphere, biosphere, anthroposphere, and cryosphere. Data come from satellites, weather stations, seismic networks, oceanographic time series, geochemical databases, and scientific text. Traditional AI models were often built for one task—land-cover classification, change detection, or a specific forecast target. Foundation models are different: they are pretrained on broad, heterogeneous data so they can be adapted to many downstream problems.
The review asks whether that flexibility can carry Earth science from perception to reasoning and, eventually, to discovery. The authors organize the field along two axes. Depth traces model capabilities from basic perception to multimodal reasoning and agentic scientific workflows. Breadth covers applications across the six major Earth spheres and their coupled processes. This is an arXiv preprint, not a new observation or a new model. Its contribution is a structured map plus a compilation of more than 200 datasets and benchmarks.
Depth: from perception to agentic workflows
The authors reviewed representative Earth foundation models and grouped them by capability stage. In the perception stage, models handle classification, segmentation, detection, and spatiotemporal prediction. Examples include weather models such as GraphCast, Pangu-Weather, and FengWu, and Earth observation systems that learn from satellite imagery. These models can be fast and accurate on their assigned tasks, but they are often limited to a single modality or objective.
The reasoning stage is where large language and multimodal models enter. Systems such as ClimaX, Aurora, and Prithvi WxC aim to learn transferable representations across weather and climate data. Multimodal models combine imagery with text, metadata, and geolocation. For example, CLLMate aligns weather and climate data with textual narratives; EarthGPT and GeoChat support visual question answering and report generation. The goal is not just to label a pixel but to answer questions about it.
The discovery stage is the newest and least mature. Agentic systems can decompose a task, choose tools, query databases, run code, and iterate. The review points to early examples including Earth-Agent, PANGAEA GPT, GISclaw, OpenEarthAgent, and ClimateAgent. Some use a single tool-augmented agent; others use multiple specialized agents for planning, data retrieval, coding, and verification. These systems are early-stage, not demonstrated autonomous scientists. But they suggest a path from one-shot prediction toward end-to-end scientific workflows.
Breadth: six spheres and coupled processes
The breadth axis is equally important. The review catalogues models and datasets for the atmosphere, hydrosphere, lithosphere, biosphere, anthroposphere, and cryosphere. Atmospheric work is the most developed, with data-driven forecasting and downscaling. Ocean and hydrology models track currents, floods, and water quality. Lithosphere models interpret seismic waves, map minerals, and estimate subsurface structure. Biosphere models estimate species ranges, canopy height, and biomass. Anthroposphere models examine urban growth, traffic, and disaster mobility. Cryosphere work is concentrated on sea ice forecasting, with glacier mapping and mass balance as related tasks.
The point is not that one model already solves all of these. It is that the same architectural ideas—pretraining, multimodal fusion, and transfer learning—are appearing across them. The survey also highlights coupled Earth system processes, where atmosphere, ocean, land, and ice interact. That coupling is precisely where isolated models struggle.
Hard limits: data, reliability, sustainability, and trust
The review is candid about what stands in the way. Earth data are heterogeneous in space, time, resolution, and meaning. A satellite image, a gridded reanalysis field, a seismic waveform, and a scientific paper do not naturally share a representation. The authors call for a unified Earth embedding—a shared numerical representation—that could encode these diverse observations without repeatedly processing raw data.
Scientific reliability is another obstacle. Models must be continually updated as new satellites and sensors come online, without forgetting what they learned before. They must also resist adversarial inputs and protect sensitive infrastructure visible in high-resolution imagery. The review raises machine unlearning (making a model forget specific data on request) and privacy as open problems.
Scalability and sustainability matter too. Training large Earth models consumes energy, and the field itself is about climate and environmental change. The authors argue for efficient adaptation, quantization (using lower-precision numbers), pruning (removing unnecessary parameters), and edge deployment. Finally, the move to embodied Earth intelligence—drones, underwater vehicles, autonomous sensing—is a future direction, not a demonstrated capability.
The most useful way to read this survey is as a map of a field in transition. Earth AI is moving from isolated predictive tools toward integrated systems that can reason with data, use scientific software, and collaborate with humans. If that transition succeeds, it could help with hazard prediction, sustainability-oriented decision-making, and planetary-scale stewardship.
But the paper does not claim that autonomous AI Earth scientists are here. The hard parts are physical consistency, uncertainty, traceability, and trust. Models that provide traceability, calibrated uncertainty, and physical consistency may be more valuable than ones that merely sound confident. The survey's emphasis is less on how many tasks a model can attempt than on whether its outputs are reliable.