Abstract illustration of an AI agent coordinating exoplanet data, spectra, and atmospheric models.

An AI Agent Ran an End-to-End Exoplanet Atmosphere Analysis. WASP-39b Showed Why Humans Still Matter.

A new preprint describes ASTER, an LLM-driven toolkit that fetched data, built spectra, and ran Bayesian retrievals for WASP-39b, while also exposing how strongly results depend on which observations and modeling choices are used.

Mar 27, 2026

WASP-39b has a radius of 1.279 Jupiter radii and a mass of 0.281 Jupiter masses. It orbits its star every 4.0553 days, with an equilibrium temperature of 1,166 K. When it passes in front of its star, its atmosphere filters starlight, leaving a spectrum with molecular fingerprints. Reading those fingerprints is a multi-step slog: query archives, download spectra, build a radiative-transfer model, then run a Bayesian retrieval—a statistical fit that estimates temperature, radius, and molecular abundances.

A new preprint asks whether an AI agent can handle that slog. In this case, the answer is yes, but not without guardrails. The paper, posted to arXiv, describes ASTER (Agentic Science Toolkit for Exoplanet Research), an orchestration framework built on the Orchestral AI framework. ASTER uses a large language model to plan and execute a chain of domain-specific tools. It is not a new telescope or a new atmospheric discovery. It is a proof of concept for making exoplanet atmosphere analysis more accessible and reproducible.

What ASTER Actually Does

ASTER connects an LLM to tools that do concrete jobs. One fetches planetary and stellar parameters from NASA's Exoplanet Archive. Another downloads public transmission spectra. A third runs TauREx, a radiative-transfer code that simulates how an exoplanet atmosphere absorbs and transmits starlight. A fourth runs TauREx retrievals using nested sampling, a Bayesian method that explores parameter space and estimates which parameters fit the data. The agent can also make plots, read files, run code, and ask the user for missing inputs.

The design matters. The LLM never executes code directly. It proposes a tool call, and the orchestration layer checks it. Orchestral includes safety hooks: heuristic rules block dangerous commands, a separate safeguard model classifies other actions, and uncertain cases go to the human for approval. It also tracks token usage and cost. The paper emphasizes that scientific judgment stays with the researcher. ASTER is meant to coordinate existing tools, not replace them.

The WASP-39b Test Drive

To test the workflow, the authors pointed ASTER at WASP-39b. The agent pulled the planet's parameters: radius 1.279 Jupiter radii, mass 0.281 Jupiter masses, equilibrium temperature 1,166 K, and stellar radius 0.939 solar radii and effective temperature 5,485 K. It generated a forward model transmission spectrum. When the required opacity files were missing, it noticed and asked for the paths. ASTER also rebinned a model spectrum to approximate JWST and Ariel instrument resolutions—without being given the exact instrument specifications. The agent later downloaded nine publicly available JWST NIRSpec observations of WASP-39b and plotted them together.

The agent then ran two independent retrievals on different public NIRSpec datasets. The first dataset spanned 3.0 to 5.5 micrometers; the second spanned 0.5 to 5.5 micrometers. The best-fit results diverged sharply. The first retrieval favored a temperature near 1,318 K and a water abundance around 10^-8.5. The second favored about 666 K and a water abundance around 10^-4.0. Planet radius stayed nearly the same, at about 1.29 to 1.30 Jupiter radii, because transit depth mostly sets that value.

Two Answers, One Important Lesson

That difference is not evidence that WASP-39b has two atmospheres. The paper explains that the datasets cover different wavelengths. The wider coverage includes more spectral features and probes different layers of the atmosphere, which can shift the best-fit temperature and abundances. The posterior distributions—the ranges of plausible parameter values—also show that the wider-coverage retrieval was more tightly constrained. The authors stress that their goal was not to improve on previous WASP-39b analyses or to settle its atmospheric properties. It was to show that an agent can autonomously orchestrate the pipeline.

That distinction is crucial. ASTER's retrievals used a reduced set of molecules and did not explore every modeling choice. The paper calls the exercise a proof of concept. It does not claim that the agent's answers are definitive. In fact, the divergent results are a useful reminder of a broader issue in exoplanet science: atmospheric retrievals depend on the data's wavelength coverage, the model's assumptions, and the molecules included.

Where the Agent Stumbles

ASTER is not flawless. The authors report that the agent sometimes needed trial and error to configure TauREx and, in some cases, hallucinated a nonexistent TauREx function. That failure helped motivate dedicated tools, which are more reliable than asking an LLM to generate code from scratch. The agent also could not fully automate downloading spectra because the archive's wget commands are not persistent. It can infer some missing details—it chose representative resolutions for JWST and Ariel bands—but the authors warn that relying on implicit knowledge increases the risk of hallucination.

There are also practical limits. ASTER currently focuses on transmission spectroscopy. Its tools can be extended by the community for broader exoplanet science. The paper is a preprint, and the WASP-39b case is a demonstration, not a new measurement.

An Assistant for the Data Deluge

Still, the bigger picture is clear. Exoplanet observations are expanding rapidly, creating a need for flexible, accessible workflows. ASTER suggests a way to wrap existing tools into an agent that can fetch data, run models, catch missing files, make plots, and compare retrievals. It lowers the barrier for newcomers and automates repetitive steps, while keeping humans in charge of interpretation.

The two WASP-39b retrievals are a fitting symbol. In this proof-of-concept case, an AI agent ran the machinery of atmospheric characterization from start to finish. But the answers it produces are only as good as the data and assumptions behind them. The useful role for ASTER may not be to give the final word on an alien atmosphere, but to help scientists manage workflows and inspect how results depend on data and assumptions—and to show, step by step, why the answers can differ. That is not a disappointment. It is a map of where human judgment still matters.

ASTER -- Agentic Science Toolkit for Exoplanet ResearchEmilie Panek, Alexander Roman, Gaurav Shukla, Leonardo Pagliaro, Katia Matcheva, Konstantin Matchevhttps://arxiv.org/abs/2603.26953v1