Submatrix and GPU-accelerated implementation of density matrix tight-binding

A Abylay Katbashev (Mulliken Center for Theoretical Chemistry, Institute for Physical and Theoretical Chemistry, University of Bonn 1 , Beringstr. 4, 53115 Bonn,) R Robert Schade (Paderborn Center for Parallel Computing, Paderborn University 2 , Warburger Str. 100, 33098 Paderborn,) M Michael Lass (dSPACE GmbH 3 , Rathenaustraße 26, 33102 Paderborn,) M Marcel Müller (Mulliken Center for Theoretical Chemistry, Clausius-Institut für Physikalische und Theoretische Chemie, Universität Bonn 1 , Beringstraße 4, 53115 Bonn,) S Stefan Grimme (Mulliken Center for Theoretical Chemistry, Clausius Institute for Physical and Theoretical Chemistry, University of Bonn, Beringstraße 4, Bonn 53115, Germany) A Andreas Hansen (Mulliken Center for Theoretical Chemistry, Clausius Institute for Physical and Theoretical Chemistry) T Thomas D. Kühne (CASUS - Center for Advanced Systems Understanding, Helmholtz-Zentrum Dresden-Rossendorf E.V. (HZDR), Untermarkt 20, Görlitz D-02826, Germany)

Abstract

Effective single-particle theories, such as Hartree–Fock, density functional theory, and tight-binding, are limited by the computational cost of the self-consistent field (SCF) procedure, which typically scales cubically with the system size. This makes large-scale applications impractical without specialized algorithms and hardware. Here, we present the submatrix and graphical processing unit (GPU)-accelerated software implementation of the PTB tight-binding potential, realized in the open-source ptb codebase [M. Mueller, A. Katbashev, and S. Ehlert (2025). “grimme-lab/ptb: v3.8.1,” Zenodo. https://zenodo.org/records/17015872]. We first benchmark a traditional diagonalization-based SCF solver against density-matrix-based purification approaches, systematically varying both system size and computer hardware. Our findings show that the usage of GPUs permits shifting the boundaries to much larger systems than previously thought feasible, achieving an overall 10–15-fold performance speedup. Second, we introduce the implementation of a decomposition-type submatrix method, specifically designed for efficient operation on mid- to large-sized systems, to address the computational overhead associated with full-system diagonalization. We demonstrate that, from a certain dimension (≈104 basis functions) on, our submatrix method reduces the overall computational cost while maintaining acceptable numerical accuracy. Our study demonstrates the significance of the interplay between modern hardware, algorithmic considerations, and novel tight-binding methods, paving the way for further development in this direction.

Article Details

Volume / Issue Vol. 163, Issue 13
Published October 07, 2025
ISSN 0021-9606
Publisher American Institute of Physics

Journal Info

The Journal of Chemical Physics

American Institute of Physics

ISSN: 0021-9606 Physical Sciences

Authors (7)

A

Abylay Katbashev

Mulliken Center for Theoretical Chemistry, Institute for Physical and Theoretical Chemistry, University of Bonn 1 , Beringstr. 4, 53115 Bonn,

R

Robert Schade

Paderborn Center for Parallel Computing, Paderborn University 2 , Warburger Str. 100, 33098 Paderborn,

M

Michael Lass

dSPACE GmbH 3 , Rathenaustraße 26, 33102 Paderborn,

M

Marcel Müller

Mulliken Center for Theoretical Chemistry, Clausius-Institut für Physikalische und Theoretische Chemie, Universität Bonn 1 , Beringstraße 4, 53115 Bonn,

S

Stefan Grimme

Mulliken Center for Theoretical Chemistry, Clausius Institute for Physical and Theoretical Chemistry, University of Bonn, Beringstraße 4, Bonn 53115, Germany

A

Andreas Hansen

Mulliken Center for Theoretical Chemistry, Clausius Institute for Physical and Theoretical Chemistry

T

Thomas D. Kühne

CASUS - Center for Advanced Systems Understanding, Helmholtz-Zentrum Dresden-Rossendorf E.V. (HZDR), Untermarkt 20, Görlitz D-02826, Germany