Protein language models for viral entry protein prediction: a multi-scale ESM2 benchmark

Viral entry proteins mediate host recognition and membrane fusion and are major targets for vaccines and antivirals. Identifying them directly from sequence data is important for characterizing emerging viruses, particularly when similarity to experimentally characterized proteins is limited. We evaluated whether frozen embeddings from the ESM2 protein language model family can support accurate binary classification of viral entry proteins. We benchmarked embeddings from three ESM2 models (650 M, 3B, and 15B parameters) with five supervised classifiers (SVM-RBF, Random Forest, XGBoost, LightGBM, and MLP) on a curated, non-redundant, balanced dataset of 1092 reviewed viral proteins from UniProt/Swiss-Prot. Sequences were clustered at 40% identity with CD-HIT and split into training ( n = 873) and independent test ( n = 219) sets. Performance improved monotonically with model scale. The best configuration, ESM2 15B with SVM-RBF, achieved 88.6% accuracy, MCC = 0.772, and ROC-AUC = 0.956 on the test set. ESM2 3B with SVM-RBF performed similarly (88.1% accuracy, MCC = 0.765, ROC-AUC = 0.953), indicating that most gains were captured by the intermediate-scale model. Classical descriptor baselines based on AAC, DPC, and PseAAC achieved lower independent test performance, with the best descriptor baseline reaching MCC = 0.537 and ROC-AUC = 0.846. In an exploratory taxid-disjoint evaluation of the ESM2 650 M plus SVM-RBF model, mean MCC and ROC-AUC were 0.734 and 0.929, respectively. We also implemented ViralEntryPred ( https://www.biochemintelli.com/viralentrypred/ ), a web application for rapid screening. Frozen ESM2 embeddings provide effective feature representations for viral entry protein prediction within a curated balanced benchmark. Larger models improved performance consistently, but with diminishing returns beyond the 3B scale, supporting intermediate-scale protein language models as a practical trade-off between accuracy and computational cost. Exploratory taxid-disjoint analysis suggested robustness to unseen taxids, but validation on completely unseen viral families remains future work.

Authors

Institutions

Publication Details

Journal
BMC Bioinformatics
Published
2026-08-25
DOI
https://doi.org/10.1186/s12859-026-06622-w
Primary Topic
Machine Learning in Bioinformatics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Protein language models for viral entry protein prediction: a multi-scale ESM2 benchmark

Lisandra Herrera Belén, Jorge F. Beltrán, Aloyma Lugo, Luis Jimenez
BMC Bioinformatics
Machine Learning in Bioinformatics
article

Protein language models for viral entry protein prediction: a multi-scale ESM2 benchmark

Lisandra Herrera Belén, Jorge F. Beltrán, Aloyma Lugo, Luis Jimenez
article en

Abstract

Viral entry proteins mediate host recognition and membrane fusion and are major targets for vaccines and antivirals. Identifying them directly from sequence data is important for characterizing emerging viruses, particularly when similarity to experimentally characterized proteins is limited. We evaluated whether frozen embeddings from the ESM2 protein language model family can support accurate binary classification of viral entry proteins. We benchmarked embeddings from three ESM2 models (650 M, 3B, and 15B parameters) with five supervised classifiers (SVM-RBF, Random Forest, XGBoost, LightGBM, and MLP) on a curated, non-redundant, balanced dataset of 1092 reviewed viral proteins from UniProt/Swiss-Prot. Sequences were clustered at 40% identity with CD-HIT and split into training ( n = 873) and independent test ( n = 219) sets. Performance improved monotonically with model scale. The best configuration, ESM2 15B with SVM-RBF, achieved 88.6% accuracy, MCC = 0.772, and ROC-AUC = 0.956 on the test set. ESM2 3B with SVM-RBF performed similarly (88.1% accuracy, MCC = 0.765, ROC-AUC = 0.953), indicating that most gains were captured by the intermediate-scale model. Classical descriptor baselines based on AAC, DPC, and PseAAC achieved lower independent test performance, with the best descriptor baseline reaching MCC = 0.537 and ROC-AUC = 0.846. In an exploratory taxid-disjoint evaluation of the ESM2 650 M plus SVM-RBF model, mean MCC and ROC-AUC were 0.734 and 0.929, respectively. We also implemented ViralEntryPred ( https://www.biochemintelli.com/viralentrypred/ ), a web application for rapid screening. Frozen ESM2 embeddings provide effective feature representations for viral entry protein prediction within a curated balanced benchmark. Larger models improved performance consistently, but with diminishing returns beyond the 3B scale, supporting intermediate-scale protein language models as a practical trade-off between accuracy and computational cost. Exploratory taxid-disjoint analysis suggested robustness to unseen taxids, but validation on completely unseen viral families remains future work.

BMC Bioinformatics
Universidad de La Frontera (CL), Universidad Católica de Temuco (CL), Universidad Santo Tomás (CL)
Openalex Percentile: Top 17%
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.