Theoretical Analysis and Optimization of Peak Memory Usage in PyPWDFT for Heterogeneous GPU High-Performance Computing

Abstract Heterogeneous high-performance computing is increasingly important for large-scale Kohn–Sham density functional theory (DFT) calculations. However, although GPU device memory provides high bandwidth, its limited capacity makes peak memory usage a critical factor that determines the maximum accessible system size. In this work, we analyze and optimize the peak memory usage of PyPWDFT, a Python native plane-wave DFT code, for heterogeneous GPU computing. The main computational stages of plane-wave DFT calculations are systematically examined to identify the dominant contributors to peak memory usage and to derive theoretical estimates of their memory requirements. Guided by this analysis, several targeted optimization strategies are introduced, including the explicit release of the CuPy memory pool, decomposition of vectorized operations, preservation of array memory contiguity, broader use of in-place operations, and additional computations when they can substantially reduce memory usage. The measured peak memory usage agrees closely with the theoretical analysis, demonstrating that with targeted memory optimization, a native Python implementation can achieve memory efficiency comparable to that of compiled-language implementations and enable DFT calculations for a 1536-atom silicon system on a single NVIDIA H200 GPU. Additional memory and performance benchmarks on A100, A800, H100, and H200 GPUs further demonstrate the general applicability of the optimized implementation across different heterogeneous GPU computing platforms.

Authors

Institutions

Publication Details

Journal
Journal of Chemical Theory and Computation
Published
2026-09-10
DOI
https://doi.org/10.1021/acs.jctc.6c01120
Primary Topic
Parallel Computing and Optimization Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Theoretical Analysis and Optimization of Peak Memory Usage in PyPWDFT for Heterogeneous GPU High-Performance Computing

Wei Hu, Jinlong Yang, Jun Gao, Bingkun Hou et al.
Journal of Chemical Theory and Computation
Parallel Computing and Optimization Techniques
article

Theoretical Analysis and Optimization of Peak Memory Usage in PyPWDFT for Heterogeneous GPU High-Performance Computing

Wei Hu, Jinlong Yang, Jun Gao, Bingkun Hou, Wenxin Peng
article en

Abstract

Abstract Heterogeneous high-performance computing is increasingly important for large-scale Kohn–Sham density functional theory (DFT) calculations. However, although GPU device memory provides high bandwidth, its limited capacity makes peak memory usage a critical factor that determines the maximum accessible system size. In this work, we analyze and optimize the peak memory usage of PyPWDFT, a Python native plane-wave DFT code, for heterogeneous GPU computing. The main computational stages of plane-wave DFT calculations are systematically examined to identify the dominant contributors to peak memory usage and to derive theoretical estimates of their memory requirements. Guided by this analysis, several targeted optimization strategies are introduced, including the explicit release of the CuPy memory pool, decomposition of vectorized operations, preservation of array memory contiguity, broader use of in-place operations, and additional computations when they can substantially reduce memory usage. The measured peak memory usage agrees closely with the theoretical analysis, demonstrating that with targeted memory optimization, a native Python implementation can achieve memory efficiency comparable to that of compiled-language implementations and enable DFT calculations for a 1536-atom silicon system on a single NVIDIA H200 GPU. Additional memory and performance benchmarks on A100, A800, H100, and H200 GPUs further demonstrate the general applicability of the optimized implementation across different heterogeneous GPU computing platforms.

Journal of Chemical Theory and Computation
University of Science and Technology of China (CN)
Openalex Percentile: Top 6%
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Theoretical Analysis and Optimization of Peak Memory Usage in PyPWDFT for Heterogeneous GPU High-Performance Computing — Wei Hu, Jinlong Yang, et al. · Journal of Chemical Theory and Computation (2026) | TGRS Research Map | TGRS