Protein language models for viral entry protein prediction: a multi-scale ESM2 benchmark
Viral entry proteins mediate host recognition and membrane fusion and are major targets for vaccines and antivirals. Identifying them directly from sequence data is important for characterizing emerging viruses, particularly when similarity to experimentally characterized proteins is limited. We evaluated whether frozen embeddings from the ESM2 protein language model family can support accurate binary classification of viral entry proteins. We benchmarked embeddings from three ESM2 models (650 M, 3B, and 15B parameters) with five supervised classifiers (SVM-RBF, Random Forest, XGBoost, LightGBM, and MLP) on a curated, non-redundant, balanced dataset of 1092 reviewed viral proteins from UniProt/Swiss-Prot. Sequences were clustered at 40% identity with CD-HIT and split into training ( n = 873) and independent test ( n = 219) sets. Performance improved monotonically with model scale. The best configuration, ESM2 15B with SVM-RBF, achieved 88.6% accuracy, MCC = 0.772, and ROC-AUC = 0.956 on the test set. ESM2 3B with SVM-RBF performed similarly (88.1% accuracy, MCC = 0.765, ROC-AUC = 0.953), indicating that most gains were captured by the intermediate-scale model. Classical descriptor baselines based on AAC, DPC, and PseAAC achieved lower independent test performance, with the best descriptor baseline reaching MCC = 0.537 and ROC-AUC = 0.846. In an exploratory taxid-disjoint evaluation of the ESM2 650 M plus SVM-RBF model, mean MCC and ROC-AUC were 0.734 and 0.929, respectively. We also implemented ViralEntryPred ( https://www.biochemintelli.com/viralentrypred/ ), a web application for rapid screening. Frozen ESM2 embeddings provide effective feature representations for viral entry protein prediction within a curated balanced benchmark. Larger models improved performance consistently, but with diminishing returns beyond the 3B scale, supporting intermediate-scale protein language models as a practical trade-off between accuracy and computational cost. Exploratory taxid-disjoint analysis suggested robustness to unseen taxids, but validation on completely unseen viral families remains future work.
Authors
- Lisandra Herrera Belén (ORCID: https://orcid.org/0000-0003-2691-0777)
- Jorge F. Beltrán (ORCID: https://orcid.org/0000-0003-2703-9629)
- Aloyma Lugo
- Luis Jimenez
Institutions
- Universidad de La Frontera (CL)
- Universidad Católica de Temuco (CL)
- Universidad Santo Tomás (CL)
Publication Details
- Journal
- BMC Bioinformatics
- Published
- 2026-08-25
- DOI
- https://doi.org/10.1186/s12859-026-06622-w
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- article
- Field-Weighted Citation Impact
- 0.00