Data‐Driven Spectral Prediction for Accelerating Large‐Scale Electronic Structure Calculations
ABSTRACT Simulating large molecular systems comprising thousands of atoms requires highly scalable methodologies. While modern Density Functional Theory (DFT) codes exhibit linear scaling, solving the associated large, sparse generalized eigenproblems remains a critical computational bottleneck on exascale architectures. In the context of the LimitX project, we propose a data‐driven framework to accelerate these calculations. By shifting the machine learning target from discrete eigenvalues to the coefficients of an interpolating Chebyshev polynomial, and by comparing both all‐atom and fragment‐based structural representations, we successfully overcome the dimensionality constraints of large‐scale spectral prediction. We investigate three machine learning models (kernel ridge regression, graph neural networks, and Random Forests) trained on a novel 2 TB dataset of protein dimers. The predicted spectra provide initial guesses that effectively bypass early self‐consistent field (SCF) iterations in BigDFT. Ultimately, these spectral predictors will be deployed to dynamically optimize upcoming rational filter‐based eigensolvers, such as FrASE, which is currently in initial development.
Authors
- Jurica Novak (ORCID: https://orcid.org/0000-0001-5744-6677)
- Xinzhe Wu (ORCID: https://orcid.org/0000-0001-5716-3116)
- Davor Davidović (ORCID: https://orcid.org/0000-0003-2649-9236)
- Gustavo Ramirez-Hidalgo
- Edoardo Di Napoli
- Abhiram Badrinarayanan
- Luigi Genovese
Institutions
- Forschungszentrum Jülich (DE)
- Commissariat à l'Énergie Atomique et aux Énergies Alternatives (FR)
- CEA Grenoble (FR)
- Institut de Recherche Interdisciplinaire de Grenoble (FR)
- Ruđer Bošković Institute (HR)
Publication Details
- Journal
- PAMM
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1002/pamm.70204
- Primary Topic
- Machine Learning in Materials Science
- Type
- article
- Field-Weighted Citation Impact
- 0.00