DyProL: Dynamic Ensemble Representation Learning for Protein–Nucleic Acid Binding Site Prediction
Protein-nucleic acid interactions play central roles in gene regulation and cellular function, and extensive efforts have been devoted to predicting nucleic acid binding sites from protein structures. However, protein-nucleic acid recognition is inherently dynamic, whereas most existing computational approaches rely on single static conformations, limiting their ability to capture conformational heterogeneity underlying binding. Here, we present DyProL, an ensemble-based conformational representation learning framework that models proteins as ensembles of conformations sampled from equilibrium-like structural distributions. DyProL learns dynamic structural features through iterative aggregation of intra- and inter-conformation geometric information, enabling representation of both local structural context and global conformational variability. Across multiple benchmarks, DyProL consistently outperforms state-of-the-art methods in nucleic acid binding site prediction, with particularly pronounced improvements under realistic settings using predicted or apo-like structures, where static methods degrade substantially. These results establish dynamic ensemble-based representations as a general and scalable paradigm for structure-based protein modeling, providing a foundation for improving a broad range of protein function prediction tasks.
Authors
- Pengpai Li (ORCID: https://orcid.org/0000-0002-2709-468X)
- Liu Rongming (ORCID: https://orcid.org/0009-0001-2681-8164)
- Yiman Liu
- Liya Liang
Institutions
- Dalian University of Technology (CN)
- Dalian University (CN)
Publication Details
- Journal
- Advanced Science
- Published
- 2026-09-03
- DOI
- https://doi.org/10.1002/advs.77501
- Primary Topic
- Protein Structure and Dynamics
- Type
- article
- Field-Weighted Citation Impact
- 0.00