PNBind : Prediction of Nucleic Acid‐Binding Sites Using Protein Structure and Protein Language Models

ABSTRACT Accurate prediction of nucleic acid‐binding sites is essential for understanding protein function and facilitating drug discovery. However, current computational methods are limited by single‐layer language model representations, Cα point clouds that discard local‐frame geometry, and feature fusion in which high‐dimensional sequence embeddings overshadow weaker evolutionary and structural signals. Here, we present PNBind, a multimodal deep learning framework that leverages protein language models, evolutionary information, and three‐dimensional structural features via a geometric vector perceptron (GVP) backbone with an invariant point attention (IPA) layer for nucleic acid‐binding site prediction. By integrating sequence representations from the ESM‐series language models with alignment‐derived evolutionary features, secondary‐structure data, and geometric features extracted from five‐atom local frames, PNBind enables effective late‐stage multimodal signal fusion. Comprehensive benchmarking on the DNA‐129, DNA‐181, RNA‐117, and RNA‐285 datasets reveals that PNBind achieves strong performance, with F 1 scores of 0.611, 0.443, 0.370, and 0.472 and MCC values of 0.586, 0.425, 0.351, and 0.416, respectively. Ablation studies confirm the dominant contribution of the protein language model backbone, while systematic evaluation of ESM‐family protein language models demonstrates that proper PLM selection significantly boosts prediction accuracy. This study demonstrates that the effective integration of protein language models and geometric deep learning offers a practical approach for accurate nucleic acid‐binding site prediction.

Authors

Institutions

Publication Details

Journal
Proteins Structure Function and Bioinformatics
Published
2026-10-08
DOI
https://doi.org/10.1002/prot.70190
Primary Topic
Machine Learning in Bioinformatics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

PNBind : Prediction of Nucleic Acid‐Binding Sites Using Protein Structure and Protein Language Models

Zhi‐Ping Liu, Bowen Shao, Pengpai Li, Yunhai Li et al.
Proteins Structure Function and Bioinformatics
Machine Learning in Bioinformatics
article

PNBind : Prediction of Nucleic Acid‐Binding Sites Using Protein Structure and Protein Language Models

Zhi‐Ping Liu, Bowen Shao, Pengpai Li, Yunhai Li, Guanghong Dang
article en

Abstract

ABSTRACT Accurate prediction of nucleic acid‐binding sites is essential for understanding protein function and facilitating drug discovery. However, current computational methods are limited by single‐layer language model representations, Cα point clouds that discard local‐frame geometry, and feature fusion in which high‐dimensional sequence embeddings overshadow weaker evolutionary and structural signals. Here, we present PNBind, a multimodal deep learning framework that leverages protein language models, evolutionary information, and three‐dimensional structural features via a geometric vector perceptron (GVP) backbone with an invariant point attention (IPA) layer for nucleic acid‐binding site prediction. By integrating sequence representations from the ESM‐series language models with alignment‐derived evolutionary features, secondary‐structure data, and geometric features extracted from five‐atom local frames, PNBind enables effective late‐stage multimodal signal fusion. Comprehensive benchmarking on the DNA‐129, DNA‐181, RNA‐117, and RNA‐285 datasets reveals that PNBind achieves strong performance, with F 1 scores of 0.611, 0.443, 0.370, and 0.472 and MCC values of 0.586, 0.425, 0.351, and 0.416, respectively. Ablation studies confirm the dominant contribution of the protein language model backbone, while systematic evaluation of ESM‐family protein language models demonstrates that proper PLM selection significantly boosts prediction accuracy. This study demonstrates that the effective integration of protein language models and geometric deep learning offers a practical approach for accurate nucleic acid‐binding site prediction.

Proteins Structure Function and Bioinformatics
Shandong University (CN), Shandong University of Political Science and Law (CN)
Openalex Percentile: Top 23%
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.