
Can One Code Run Black Hole Simulations on Both CPUs and GPUs?
KHARMA, a new open-source GRMHD code, is designed to simulate black hole accretion efficiently on CPUs and GPUs, and its modular packages open the door to longer, more physics-rich Event Horizon Telescope comparisons.
Aug 2, 2024
To turn a blurry ring of light around a supermassive black hole into a physical story, astronomers need more than telescopes. They need simulations: computer models of magnetized plasma swirling through curved spacetime. In a new preprint chapter, Ben S. Prather of Los Alamos National Laboratory describes KHARMA, an open-source code built to run those simulations efficiently on both conventional processors and graphics processors, which are increasingly central to supercomputing.
KHARMA Is Already Feeding Event Horizon Telescope Comparisons
The Event Horizon Telescope made history by imaging the supermassive black holes at the centers of M87 and the Milky Way. Those images are compared with models of how plasma falls onto a black hole, heats up, and radiates. General-relativistic magnetohydrodynamics, or GRMHD, is the tool for that job: it tracks magnetized fluid in curved spacetime around a black hole. Because the equations are so demanding, researchers often run libraries of simulations, varying spin, magnetic field strength, and other unknowns, then compare predicted images with telescope data.
KHARMA is already one of the primary GRMHD codes used by the Event Horizon Telescope Collaboration. The paper says it furnished many of the GRMHD simulations used in official Sagittarius A* results and collaboration projects. That matters because Sagittarius A* varies quickly: to model just a few days of its behavior, simulations must run long enough to forget their starting conditions.
One C++ Code, Kokkos, Parthenon, and Run-Time Packages
The central design choice behind KHARMA is performance portability: the same C++ source code can compile and run efficiently on traditional CPUs and on GPUs from major vendors. It uses the Kokkos programming model, which turns abstract parallel loops into machine-specific code at compile time, and the Parthenon framework for adaptive mesh refinement. Adaptive mesh refinement, or AMR, puts more grid cells where detail is needed and fewer where it is not. That helps with a geometric problem: in spherical coordinates, cells bunch up near the poles, forcing impractically small time steps. KHARMA was designed with AMR in mind from the outset; after face-centered constrained magnetic field transport was added, it could support refined meshes for magnetized runs too.
The code is also modular. Its functions are separated into packages, each representing an algorithmic component or a piece of physics. Algorithmic choices and magnetic field transport can be swapped at run time. All features and problem setups are available at run time, so a single compiled binary can run any implemented problem, a design choice useful for testing and for emulating other codes' schemes.
Benchmarks, Extended Physics, and Tests Show the Trade-Offs
The paper reports benchmarks on several supercomputers, including Summit, Frontier, Delta, and Longhorn. Running a standard torus accretion problem in ideal GRMHD, KHARMA reached about 21 million zone-cycles per second on an Nvidia V100 GPU and about 35 million on an A100; a zone-cycle here means one grid cell advanced one time step. On one Delta node with four A100s, the code performed like more than 50 nodes of the Stampede2 CPU system for a production-size problem.
Frontier, an AMD MI250X machine, was a different story. KHARMA reached about 100 million zone-cycles per second per node there, but register pressure in its complex kernels limited thread occupancy. That made a Frontier node roughly comparable to a Summit node, even though Frontier has more memory bandwidth and compute. The paper suggests performance could improve by a factor of two if the code were fully limited by memory bandwidth rather than computation. Those numbers are benchmarks, not astrophysical measurements, and performance varies with the physics enabled: more advanced features can run 10 to 30 percent slower than the base configuration. Extended GRMHD is far costlier, about five times slower on CPUs and 10 to 20 times slower on GPUs.
Ideal GRMHD omits effects such as viscosity, heat conduction, and resistivity that can matter in real accretion flows. KHARMA includes packages for electron temperature evolution, viscous hydrodynamics, and an Extended GRMHD treatment with pressure anisotropy and heat conduction. That last extension is aimed at low-collision, low-accretion systems such as Sagittarius A*, where the paper hopes plasma physics can help explain rapid variability and EHT comparisons. The code also supports chained multi-scale simulations: in one ongoing effort, eight overlapping annuli span scales from the event horizon out to galactic distances, each feeding the next, with the goal of improving the simple sub-grid models that large galactic simulations use for supermassive black hole feedback. Radiative effects, important at high accretion rates, are being developed but are not presented as part of this chapter's validation and performance results.
The paper emphasizes validation. After every code commit, regression tests run automatically on CPUs and manually on GPUs to check convergence on GRMHD and Extended GRMHD problems. At the time of writing, 46 ideal MHD linear-mode convergence tests were performed, covering different sub-step types, integrators, magnetic field transports, and mesh refinements. The L1 norm, a grid-summed measure of difference from the analytic solution, decreases with resolution as expected for a second-order scheme; the paper reports powers between -1.9 and -2.1. Other tests check restarting from checkpoints, determinism, and mesh-refined versions. KHARMA implements the HARM scheme and was designed first to match the earlier iharm3D code, which was validated in the Event Horizon Telescope GRMHD code comparison project. Where KHARMA uses newer components, the paper argues they are tested as drop-in replacements. Still, the paper notes that no separate unit tests exist, and that debugging interactions between packages remains a challenge; the long-term success of its readability strategy remains to be seen.
KHARMA is not a discovery about a particular black hole. It is infrastructure: a faster, more flexible way to ask what the Event Horizon Telescope's rings and polarization patterns might mean. The paper does not present any one simulation as final; its argument is that performance portability and modular physics let researchers run more models, longer, with fewer compromises. That matters as telescope data improve and as computing shifts toward specialized accelerators. If one code can move across those machines without being rewritten, black hole simulations can keep pace. The reward is not just prettier images. It is a better chance to turn a ring of light into a tested story about gravity, plasma, and the edge of a black hole.