
A New Code Simulates 232 Billion Dark-Matter Particles—and Barely Slows Down at 512 Nodes
CUBE2, a new open-source N-body code, ran two 6144³-particle cosmological simulations on 512 nodes and reported 94% weak-scaling efficiency, with matter clustering matching CAMB predictions.
Dec 14, 2025
To understand how dark matter built the cosmic web, cosmologists run simulated universes. The rules are simple: mass attracts mass. The bookkeeping is not. Every simulated particle needs a position and a velocity, and next-generation galaxy surveys require simulations with enormous volumes and high mass resolution, pushing toward trillion-particle runs. At that scale, memory can stop a simulation before gravity even gets started.
A new code called CUBE2, described in a preprint on arXiv, attacks that bottleneck. It is a methods paper: not a new observation of the sky, but a new way to build simulated skies relevant to next-generation surveys such as DESI, LSST, and China's CSST. The authors report that CUBE2 ran two cosmological simulations with 6144³ particles—about 232 billion—on 512 computing nodes, while keeping memory overhead low and scaling close to ideal.
A Simulated Universe in 6 Bytes per Particle
An N-body simulation represents matter as many discrete particles and lets gravity move them. It is a workhorse for modeling how cold dark matter grew into the large-scale structure that galaxies trace. The hard part is force calculation. Directly computing every pair of particles scales as N squared, impossible for huge N, so codes approximate.
CUBE2 combines two standard ideas. A Particle-Mesh (PM) method assigns particles to a grid and solves gravity with fast Fourier transforms, a mathematical technique for handling fields efficiently. Particle-Particle (PP) calculations handle close pairs directly. CUBE2 uses three PM layers: a coarse global mesh for long-range gravity, a local fixed mesh for intermediate scales, and an adaptive finer mesh for dense regions. PP then cleans up the shortest ranges. An optimized Green's function, essentially a mathematical rule for how mass pulls on its surroundings, bridges the layers.
Memory is the other constraint. Traditional floating-point storage can cost about 24 bytes per particle for positions and velocities. CUBE2 uses a fixed-point format: coordinates are stored as small integers relative to a grid, with separate bulk fields, and particle ordering carries additional information. The paper says this can reduce the basic cost toward 6 bytes per particle, or 12 bytes when higher precision is needed.
512 Nodes, 94% Weak Scaling, 28.1× Speedup
The tests ran on the Advanced Computing East China Sub-center, using 512 nodes and 16,384 CPU cores (32 cores per node). Each simulation followed 6144³ particles in a flat ΛCDM cosmology—cold dark matter plus dark energy. The two boxes were 2400 and 1200 Mpc h⁻¹ across, roughly 11.2 and 5.6 billion light-years for h = 0.70. Initial conditions were set at redshift 200, an early epoch, using the Zel'dovich approximation. Each run produced 100 snapshots, and the total computation times were 23.8 and 40.1 days.
These are simulated universes, not telescope observations. Their value is as a numerical laboratory: they let researchers connect early conditions to the large-scale structure seen today, and they can later serve as starting points for mock galaxy catalogs.
The headline result is scalability. In a weak-scaling test, where the problem size grows in proportion to the number of computers, the 512-node run took only 6.8% more time per timestep than a single-node problem. That is 94% weak scaling. In a strong-scaling test, where the problem stays fixed and more cores are added, 32 cores ran 28.1 times faster than one core, or 88% strong scaling.
The code also balances uneven work. Dense cosmic regions take longer to simulate than voids. CUBE2 sorts the smallest computational tiles by density and gives the heaviest ones to parallel teams first. The paper reports that this reduced load imbalance by 11.4% and reached near-perfect balancing at 99.6%. Memory overhead from clustering stayed below 10% in the larger box and below 30% in the smaller one.
Matching CAMB—and What the Code Still Leaves Out
CUBE2's force calculation was tested on random particle pairs by comparing the summed PM and PP forces with a reference force. The errors appear at the expected force-matching scales, where the code deliberately softens gravity to avoid nonphysical particle scattering.
The final snapshot at redshift zero, the present epoch, was used to compute the matter power spectrum, a statistical measure of clustering across scales. Both simulations align well with CAMB theoretical predictions for linear and nonlinear evolution. That agreement is a consistency check, not an independent measurement of the universe.
This is a code paper. It does not announce a new cosmological discovery. The simulations assume cold dark matter and dark energy, use Newtonian gravity for the simulated scales, and neglect small-scale baryonic effects at late epochs. The initial conditions use first-order Lagrangian perturbation theory at high redshift, a standard but simplified choice. The paper says detailed scientific analysis will come later. CUBE2 is also configured with neutrino modules, but these performance runs use cold dark matter and dark energy. The manuscript is a preprint on arXiv.
Next-generation surveys will map enormous volumes and faint galaxies, testing dark energy, neutrino mass, and primordial non-Gaussianity. Those analyses need simulations that are both large and accurate. CUBE2's bet is that clever storage, multi-level force calculation, and load balancing can make trillion-particle simulations feasible on modest platforms.
The universe is not getting smaller. The simulations are getting smarter—byte by byte, grid by grid.