Abstract illustration of a gravitational-wave chirp passing through separate frequency channels in a neural network.

Can a Smaller Neural Network Clean Gravitational Waves Better by Listening in Bands?

A controlled comparison of five neural-network architectures finds that a frequency-aware design reconstructs simulated and real black hole merger signals better than a larger U-Net, while leaving glitches as the main obstacle.

Sep 7, 2026

A gravitational-wave signal from two merging black holes is a chirp: a rising tone that sweeps from a low rumble to a final thud. By the time it reaches Earth, that thud is buried in detector noise. Recovering the waveform is not cosmetic. It is the step that lets astronomers estimate black hole masses and spins, test general relativity in strong gravity, and study how binary black holes form.

A preprint posted to arXiv by Rohan Raha and Prayush Kumar compares five neural-network architectures for that cleaning job. The authors describe it as the first controlled comparison of neural-network architectures for gravitational-wave denoising, trained identically across the full astrophysically motivated spinning binary-black-hole parameter space. The result is not simply that bigger is better. A network built around the frequency structure of a coalescence outperformed the larger U-Net.

Why denoising is a post-detection problem

Gravitational-wave denoising is usually applied after a candidate signal has already been found. The search itself may come from matched filtering, which compares detector data against a bank of predicted waveforms. That method is powerful but computationally expensive. As next-generation detectors push event rates higher, the cost of matched filtering could become prohibitive.

Deep learning offers a different path: a trained network can reconstruct a waveform in real time. But many existing denoisers were developed on narrow parameter spaces, often with non-spinning black holes, limited mass ratios, or short signal windows. That makes principled comparison and reliable deployment difficult.

This work tests a harder, more realistic population. The training and testing datasets each contain 20,000 simulated binary-black-hole signals, each two seconds long at a sampling rate of 4096 Hz. Component masses span 5 to 100 solar masses, mass ratios go down to 1/6, and spins reach 0.99. Half the systems are precessing, meaning their spins are misaligned with the orbital axis, so the binary's orbital plane wobbles; a quarter are aligned-spin, and a quarter are non-spinning. The waveforms are injected into simulated colored Gaussian noise matching Advanced LIGO design sensitivity and projected onto the Hanford detector.

Five networks, one controlled test

The authors compare a baseline recurrent network, a convolutional-recurrent hybrid, a transformer-style architecture called ACRED-Net, a large U-Net with attention gates, and their own Multi-Scale Frequency-Aware design. All five see the same data, use the same loss function and optimizer, and stop training by the same rule. To keep the networks from hallucinating signals in pure noise, the training set includes 30 percent signal-free samples, a fraction chosen through an ablation study.

The Multi-Frequency architecture is the standout. Instead of forcing one network to handle every part of the chirp, it routes the input through seven parallel branches. Three large-kernel branches specialize in the low-frequency inspiral. Four small-kernel branches target the higher-frequency merger and ringdown. A cross-frequency integration stage then combines what the branches learn.

That design achieved the best validation reconstruction fidelity of the five. Its best validation mean squared error, a measure of how far the reconstructed waveform strays from the true one, was 1.15 times 10 to the minus 3. That was 8.7 percent better than the much larger U-Net, 16.1 percent better than ACRED-Net, 39.8 percent better than the CNN-LSTM hybrid, and 70.5 percent better than the baseline recurrent network. Crucially, it did this with about 59 million parameters, roughly one-third as many as the 172-million-parameter U-Net.

The authors argue that the gain comes from matching the network structure to the spectral anatomy of a coalescence: inspiral, merger, and ringdown. Brute-force scaling, in this test, was less effective than a physically motivated decomposition.

What the network recovers — and what it misses

Across a population of simulated signals, the Multi-Frequency model performed well. Depending on spin category, 55 to 64 percent of test events achieved a mismatch below 0.01, and 80 to 90 percent fell below 0.03. Mismatch here is one minus the noise-weighted overlap with the true waveform, so lower is better. Precessing systems remained the hardest, as expected: their wobbling orbital plane imprints quasi-periodic amplitude and phase modulations that are harder to reconstruct than the smoother non-spinning or aligned-spin signals.

The team also built population-level uncertainty bands from the residuals of an independent evaluation ensemble. These are not per-event Bayesian posteriors. They describe the typical residual size for a given spin category at a fixed signal-to-noise ratio of 20. The authors then checked their statistical calibration with a probability-probability diagnostic, which asks whether a claimed n-sigma band really contains the true residual an n-sigma fraction of the time. The bands passed that test to within a few percent.

From simulated noise to real detectors

The more striking test came when the network, trained only on simulated Hanford noise, was applied without retraining to real LIGO-Virgo-KAGRA data. The authors examined three confirmed binary-black-hole events spanning three observing runs: GW150914 from O1, GW200129 from O3b, and GW231226 from O4. The network recovered merger morphology in both the Hanford and Livingston detectors.

Against maximum-likelihood waveform templates, the denoised outputs achieved Pearson correlations above 0.95 and energy-normalized mean squared errors below 0.13. The Livingston data for GW200129 included glitch mitigation, an important caveat because a noise transient overlapped the signal. But the overall consistency suggests the network learned signal features that transfer across detector noise environments, not just Hanford-specific artifacts.

When there is no signal

The authors also ran the network over 24 hours of real Hanford noise from the O4a run, excluding known events. After processing, 43,083 two-second segments remained, corresponding to about 12 hours of analyzed background over the same 1-second analysis window. For most segments, the network suppressed its output almost completely: the median output norm was 0.046, close to the ideal of zero.

Rare high outputs did occur. The authors traced them almost exclusively to real non-Gaussian noise transients, or glitches. Because the network was trained only on stationary simulated Gaussian noise, it has no learned basis for suppressing genuine glitches. In those cases, it removed the surrounding background and left the transient's structure largely intact, effectively unveiling the glitch rather than inventing a signal.

That is a promising sign for signal-versus-noise discrimination, but it is not a detection study. The paper frames denoising as post-detection reconstruction, and the authors say a dedicated detection study is future work.

What could still go wrong

Several limitations keep this from being a turnkey pipeline. The model was trained only on simulated Hanford noise and applied to each detector independently, not through a coherent multi-detector reconstruction. Separate Hanford and Livingston outputs are therefore not guaranteed to correspond to consistent source parameters. Glitches remain the dominant failure mode. The parameter space is restricted to quasi-circular binary black holes, leaving out neutron-star systems, eccentric orbits, and intermediate-mass-ratio inspirals.

Out-of-distribution tests also show where the method bends. On extreme mass ratios beyond the training range, only 36 percent of injections achieved mismatch below 0.01, compared with 55 to 64 percent for the trained categories. High signal-to-noise events still performed much better, and a spin distribution resembling hierarchical-merger remnants degraded only modestly.

The practical contribution is a reproducible benchmark: the authors released trained model weights and architecture definitions. The broader lesson is that for gravitational-wave denoising, structure may matter as much as scale. The chirp is still buried in noise, but a network that knows where to listen in frequency is learning to pull it out.