Implementation of the Multigrid Gaussian-Plane-Wave Algorithm with GPU Acceleration in PySCF

Abstract We introduce a GPU-accelerated multigrid Gaussian-Plane-Wave density fitting (FFTDF) approach for efficient Fock builds and nuclear gradient evaluations within Kohn–Sham density functional theory, as implemented in the GPU4PySCF module of PySCF. Our CUDA kernels employ a grid-based parallelization strategy for contracting Gaussian basis function pairs and achieve 50% to 80% of the FP64 peak performance on NVIDIA GPUs, with no loss of efficiency for high angular momentum (up to f-shell) functions. Benchmark calculations on molecules and solids with up to 1536 atoms and 20480 basis functions show up to 25 × speedup on an H100 GPU relative to the CPU implementation on a 28-core shared-memory node. For a 256-water cluster, the ground-state energy and nuclear gradients can be computed in ∼30 seconds on a single H100 GPU. This implementation serves as an open-source foundation for many applications, such as ab initio molecular dynamics and high-throughput calculations.

Authors

Institutions

Publication Details

Journal
Journal of Chemical Theory and Computation
Published
2026-09-17
DOI
https://doi.org/10.1021/acs.jctc.6c00947
Primary Topic
Advanced Chemical Physics Studies
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Implementation of the Multigrid Gaussian-Plane-Wave Algorithm with GPU Acceleration in PySCF

Qiming Sun, Garnet Kin-Lic Chan, Junjie Yang, Xing Zhang et al.
Journal of Chemical Theory and Computation
Advanced Chemical Physics Studies
article

Implementation of the Multigrid Gaussian-Plane-Wave Algorithm with GPU Acceleration in PySCF

Qiming Sun, Garnet Kin-Lic Chan, Junjie Yang, Xing Zhang, Rui Li, Yuanheng Wang
article en

Abstract

Abstract We introduce a GPU-accelerated multigrid Gaussian-Plane-Wave density fitting (FFTDF) approach for efficient Fock builds and nuclear gradient evaluations within Kohn–Sham density functional theory, as implemented in the GPU4PySCF module of PySCF. Our CUDA kernels employ a grid-based parallelization strategy for contracting Gaussian basis function pairs and achieve 50% to 80% of the FP64 peak performance on NVIDIA GPUs, with no loss of efficiency for high angular momentum (up to f-shell) functions. Benchmark calculations on molecules and solids with up to 1536 atoms and 20480 basis functions show up to 25 × speedup on an H100 GPU relative to the CPU implementation on a 28-core shared-memory node. For a 256-water cluster, the ground-state energy and nuclear gradients can be computed in ∼30 seconds on a single H100 GPU. This implementation serves as an open-source foundation for many applications, such as ab initio molecular dynamics and high-throughput calculations.

Journal of Chemical Theory and Computation
California Institute of Technology (US), University of Siedlce (PL)
National Energy Research Scientific Computing Center, Office of Science, Lawrence Berkeley National Laboratory, Energy Frontier Research Centers
Affordable and clean energy
Openalex Percentile: Top 74%
Advanced Chemical Physics Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.