
AI Can Learn Which Galaxy Features Are Real—and Which Belong to the Telescope
A new deep-learning framework uses overlapping galaxy images from the DESI Legacy Imaging Surveys and Hyper Suprime-Cam to separate intrinsic structure from instrument artifacts, enabling counterfactual views and instrument-independent searches.
Apr 10, 2026
Imagine two photographs of the same galaxy. In one, a faint spiral arm is crisp. In the other, it smears into a soft blob. The galaxy did not change. The telescope did.
That is everyday multi-instrument astronomy. A telescope records photons after they pass through optics, detector characteristics, the atmosphere, observing conditions, and calibration. The observation mixes the intrinsic physical signal, instrument-specific artifacts, and noise. A new preprint proposes a way to separate them.
One Galaxy, Two Telescopes, Two Different Pictures
Astronomers have long used multiple instruments to tell real signals from artifacts. Gravitational-wave detectors, for example, need coincident observations to confirm a faint event. The new work applies that logic to galaxy images, using the DESI Legacy Imaging Surveys and the Hyper Suprime-Cam (HSC) Survey.
Legacy is wide and shallow, covering about 20,000 square degrees. HSC is deeper and sharper but covers only about 1,200 square degrees. The same galaxy can look remarkably different in each. The authors train on roughly 100,000 cross-matched galaxy images from the two surveys.
A Training Trick: Hide the Answer, Then Reconstruct It
The framework uses two encoders. One sees the target galaxy through a different instrument and is pushed to capture its intrinsic physical properties. The other sees different galaxies through the target instrument and is pushed to capture the instrument's distortions. A decoder combines those learned representations to reconstruct an unseen anchor image.
Crucially, the anchor image is never fed to the encoders. It is only the target. That counterfactual generation objective forces the physics code to contain what is shared across instruments and the instrument code to contain what is specific to the measurement. The decoder is a flow-matching model, a generative method that learns a distribution of possible images rather than a single best guess.
What the Model Learned to Separate
Testing the learned internal representations, or latent spaces, a two-dimensional projection called UMAP showed that cross-survey pairs of the same galaxy land near each other in physics space. In instrument space, HSC and Legacy observations form distinct clusters. The model was not trained with a contrastive loss, yet the physics representations aligned across surveys.
Downstream regressions found that physics latents predict galaxy properties—redshift, stellar mass, half-light radius, ellipticity, specific star formation rate, metallicity, and mass-weighted stellar age—at a level comparable to AION-1, a large astronomical foundation model that treats surveys as separate modalities. Instrument latents were better at predicting observing conditions such as point-spread function size (the blur from atmosphere and optics), exposure count, and survey depth. The key diagnostic is asymmetry: physics latents held less instrument information than a cross-prediction baseline (a model trained to predict one survey's instrument properties from the other survey's images), suggesting the architecture actively erased instrument-related signals beyond what sky position alone would explain.
A Legacy Image That Behaves Like a Hyper Suprime-Cam Image
The generative decoder can produce counterfactual views. Given a Legacy image, it predicts how the galaxy would appear to HSC. Given an HSC image, it predicts a Legacy-like view. On 256 held-out galaxies, the generated posterior samples had a mean squared error of 0.081 for HSC-anchored reconstructions (targeting HSC from Legacy conditioning) and 0.197 for Legacy-anchored ones (targeting Legacy from HSC conditioning), computed on preprocessed, standardized pixel values.
In a more concrete test, a ResNet—a standard image neural network—trained on 25,000 real HSC images to measure galaxy ellipticity was applied without retraining to generated HSC-like images from Legacy data. The predictions matched ground truth about as well for generated images (R² of 0.81) as for real HSC images (R² of 0.82), where 1 would be perfect. For this ellipticity-based morphology test, that suggests an existing HSC analysis pipeline could be used on Legacy data after passing it through the model.
The embeddings also support instrument-independent similarity search. In physics space, a query returns physically similar galaxies from either survey. In instrument space, it returns objects with similar noise and instrument conditions. That could help hunt rare objects, such as strong gravitational lenses, by first predicting which Legacy galaxies are worth a deeper look.
Where the Method Still Struggles
The method requires overlapping observations, limiting it to sky seen by both instruments. It also favors shared information: the objective discards details unique to a higher-quality instrument, such as HSC's ability to resolve morphology that Legacy cannot. The authors propose future per-instrument residual latents to preserve that extra information.
The uncertainty calibration is not perfect. Posterior samples underestimate the true pixel-wise variance by about 15%, and the model shows mild overconfidence, especially when generating higher-resolution HSC-like images. The instrument latent space still retains some physical information, such as redshift, stellar mass, and morphology, because those properties are entangled with observing conditions. The paper is a preprint, posted to arXiv on April 10, 2026, and has not been peer-reviewed. The galaxy demonstration uses a cross-matched sample and held-out tests; it is not a full survey validation.
A Recipe for Noisy Data Beyond Galaxies
The broader recipe is general: build training pairs from overlapping observations, treat sensor and modality effects as augmentations, and learn invariant representations through counterfactual generation. The authors plan to apply it to tens of millions of light curves—brightness measurements over time—from NASA's TESS and Kepler missions. The goal is not to replace real observations but to reduce the search space for follow-up.
The same galaxy can look remarkably different through two telescopes. This work aims to let astronomers see the galaxy, not just the instrument's accent.