Capturing Spatial Covariance for Efficient Multipolar Electrostatic Parameterization in DNA via Multioutput Learning

Abstract In DNA force field development, introducing high-rank atomic multipole moments (AMMs) to describe electron cloud anisotropy improves the accuracy of electrostatic energy calculations. However, high-rank AMMs have many components. Existing machine learning methods usually model each component independently, which increases model size and computational cost. This approach ignores the correlations among components and cannot balance accuracy and efficiency. In this study, a DNA molecular electrostatic energy data set was constructed for six base pairing types to obtain AMMs labels up to the hexadecapole moment. We evaluated two task grouping strategies alongside three multioutput Gaussian process regression (MOGPR) architectures (the Linear Model of Coregionalization (LMC), Intrinsic Coregionalization Model (ICM), and Semiparametric Latent Factor Model (SLFM)). Results showed that the Component strategy, which groups components by rank, reduced the total number of models by 80% for each data set. This grouping lowered the computational and maintenance costs of force field construction while reducing interference between components of different ranks. Among the MOGPR architectures, SLFM balanced accuracy and efficiency. It described the spatial covariance relationships among tasks through a low-rank structure, overcoming the limitations of independent prediction. In molecular electrostatic energy reconstruction, SLFM showed an even greater advantage. On the CG data set, which has high conformational dispersion, the mean absolute error of SLFM decreased by 46.5% compared with the single-task method. This study combined quantum chemical topology and multitask learning, providing a reference for predicting high-rank AMMs and developing DNA force fields.

Authors

Institutions

Publication Details

Journal
The Journal of Physical Chemistry B
Published
2026-09-22
DOI
https://doi.org/10.1021/acs.jpcb.6c04920
Primary Topic
Protein Structure and Dynamics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Capturing Spatial Covariance for Efficient Multipolar Electrostatic Parameterization in DNA via Multioutput Learning

Yongna Yuan, Ruisheng Zhang, Wei Su, Zhenyu Liu
The Journal of Physical Chemistry B
Protein Structure and Dynamics
article

Capturing Spatial Covariance for Efficient Multipolar Electrostatic Parameterization in DNA via Multioutput Learning

Yongna Yuan, Ruisheng Zhang, Wei Su, Zhenyu Liu
article en

Abstract

Abstract In DNA force field development, introducing high-rank atomic multipole moments (AMMs) to describe electron cloud anisotropy improves the accuracy of electrostatic energy calculations. However, high-rank AMMs have many components. Existing machine learning methods usually model each component independently, which increases model size and computational cost. This approach ignores the correlations among components and cannot balance accuracy and efficiency. In this study, a DNA molecular electrostatic energy data set was constructed for six base pairing types to obtain AMMs labels up to the hexadecapole moment. We evaluated two task grouping strategies alongside three multioutput Gaussian process regression (MOGPR) architectures (the Linear Model of Coregionalization (LMC), Intrinsic Coregionalization Model (ICM), and Semiparametric Latent Factor Model (SLFM)). Results showed that the Component strategy, which groups components by rank, reduced the total number of models by 80% for each data set. This grouping lowered the computational and maintenance costs of force field construction while reducing interference between components of different ranks. Among the MOGPR architectures, SLFM balanced accuracy and efficiency. It described the spatial covariance relationships among tasks through a low-rank structure, overcoming the limitations of independent prediction. In molecular electrostatic energy reconstruction, SLFM showed an even greater advantage. On the CG data set, which has high conformational dispersion, the mean absolute error of SLFM decreased by 46.5% compared with the single-task method. This study combined quantum chemical topology and multitask learning, providing a reference for predicting high-rank AMMs and developing DNA force fields.

The Journal of Physical Chemistry B
Lanzhou University of Technology (CN), Lanzhou University (CN)
Affordable and clean energy
Openalex Percentile: Top 18%
Protein Structure and Dynamics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.