Load-balanced diffusion Monte Carlo method with lattice regularization

K Kousuke Nakano (Center for Basic Research on Materials (CBRM), National Institute for Materials Science (NIMS) 1 , 1-2-1 Sengen, Tsukuba, Ibaraki 305-0047,) S Sandro Sorella (International School for Advanced Studies (SISSA) 2 , Via Bonomea 265, 34136 Trieste,) M Michele Casula (Institut de Minéralogie, de Physique des Matériaux et de Cosmochimie)

Abstract

Ab initio quantum Monte Carlo (QMC) is a stochastic approach for solving the many-body Schrödinger equation without resorting to one-body approximations. QMC algorithms are readily parallelizable via ensembles of Nw walkers, making them well suited to large-scale high-performance computing. Among the QMC techniques, diffusion Monte Carlo (DMC) is widely regarded as the most reliable since it provides the projection onto the ground state of a given Hamiltonian under the fixed-node approximation. One practical realization of DMC is the lattice regularized diffusion Monte Carlo (LRDMC) method, which discretizes the Hamiltonian within the Green’s function Monte Carlo framework. DMC methods—including LRDMC—employ the so-called branching technique to stabilize walker weights and populations. At the branching step, walkers must be synchronized globally; any imbalance in per-walker workload can leave central processing unit (CPU) or graphics processing unit (GPU) cores idle, thereby degrading overall hardware utilization. The conventional LRDMC algorithm intrinsically suffers from such load imbalance, which grows as log(Nw), rendering it less efficient on modern parallel architectures. In this work, we present an LRDMC algorithm that inherently addresses the load imbalance issue and achieves significantly improved weak-scaling parallel efficiency. Using the binding energy calculation of a water–methane complex as a test case, we demonstrated that the conventional and load-balanced LRDMC algorithms yield consistent results. Furthermore, by utilizing the Leonardo supercomputer equipped with NVIDIA A100 GPUs, we demonstrated that the load-balanced LRDMC algorithm can maintain extremely high parallel efficiency (∼98%) up to 512 GPUs (corresponding to Nw = 51 200), together with a speedup of ×1.24 if directly compared with the conventional LRDMC algorithm with the same number of walkers. The speedup stays sizable, i.e., × 1.18, even if the number of walkers is reduced to Nw = 400.

Article Details

Volume / Issue Vol. 163, Issue 19
Published November 21, 2025
ISSN 0021-9606
Publisher American Institute of Physics

Journal Info

The Journal of Chemical Physics

American Institute of Physics

ISSN: 0021-9606 Physical Sciences

Authors (3)

K

Kousuke Nakano

Center for Basic Research on Materials (CBRM), National Institute for Materials Science (NIMS) 1 , 1-2-1 Sengen, Tsukuba, Ibaraki 305-0047,

S

Sandro Sorella

International School for Advanced Studies (SISSA) 2 , Via Bonomea 265, 34136 Trieste,

M

Michele Casula

Institut de Minéralogie, de Physique des Matériaux et de Cosmochimie