Protein–Protein Interface Residue Prediction by Integrating Protein Language Representations and Structural Graph Attention
Abstract Accurate residue-level prediction of protein–protein interaction interfaces is important for elucidating molecular recognition mechanisms and guiding targeted drug design. However, interface residues are relatively sparse and governed by complex spatial dependencies, making their accurate identification challenging. In this study, we developed EvoStruct-GAT, a multimodal graph attention model based on the PDBbind v2020.R1 data set. To reduce information leakage arising from sequence homology, receptor and ligand sequences were clustered independently, and the resulting clusters were used to construct training, validation, and test sets with no overlap in either receptor or ligand clusters. EvoStruct-GAT represents the target protein as a residue-level graph and integrates ESM-2 protein language model representations, DSSP-derived structural descriptors, and distance-aware graph attention, together with a combined strategy for addressing class imbalance. Using a minimum cross-partner heavy-atom distance of less than 5 Å to define interface residues, the model achieved an AUROC of 0.8081, an AUPRC of 0.4475, and an F1 score of 0.4562 on the independent test set. Alternative labels based on 4.0 Å, 4.5 Å, and ΔSASA definitions produced comparable AUROC values, indicating that the residue-ranking capability was reasonably robust across the tested interface definitions. Ablation experiments identified ESM-2 representations as the principal contributor to performance, with DSSP information and inter-residue message passing providing additional value. Under a common evaluation protocol, EvoStruct-GAT outperformed simplified internal baselines and showed complementary performance characteristics in external comparisons with publicly available models. These results demonstrate that EvoStruct-GAT can effectively rank potential interface residues without requiring information about a known binding partner, thereby supporting candidate-site selection, experimental validation, and downstream structural and functional analyses.
Authors
- Wei He (ORCID: https://orcid.org/0000-0002-9962-4627)
- Zhen Hou (ORCID: https://orcid.org/0000-0002-5476-1441)
- Hongquan Li (ORCID: https://orcid.org/0000-0002-2030-5587)
- Jiarui Li
- Weinan Cao (ORCID: https://orcid.org/0009-0007-9595-5588)
- Shuo Sun
- Hua Yang
Institutions
- Shanxi Agricultural University (CN)
- Shanxi University (CN)
Publication Details
- Journal
- ACS Omega
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1021/acsomega.6c08819
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- article
- Field-Weighted Citation Impact
- 0.00