PrePssmCas: fusing pre-trained features and PSSM features for Cas protein classification
Abstract Background Clustered regularly interspaced short palindromic repeats (CRISPR) and their associated (Cas) proteins play a vital role in prokaryotic adaptive immune systems, defending against foreign genetic material. However, the classification of CRISPR-Cas systems remains challenging due to limited direct identification techniques. Distinct classes of Cas proteins perform different roles, offering an opportunity for system classification through their identification. The accurate identification of specific Cas proteins is therefore essential for advancing CRISPR-Cas system classification. Results To enhance Cas protein discrimination, two complementary types of features were extracted from protein sequences: pre-trained protein language model embeddings and fixed-dimensional features derived from position-specific scoring matrix (PSSM) profiles. Through systematic comparison of five pre-trained models (Prot_BERT, ALBERT, ProtXLNet, ProtT5, and ESM1b) and six PSSM variants, combined with an attention-based aggregation strategy, ESM1b features combined with RPM-PSSM features demonstrated the most comprehensive representation of Cas proteins. By employing random forest-based feature selection, a 143-dimensional feature vector was obtained, comprising 87 pre-trained features and 56 RPM-PSSM features. The selected features achieved an accuracy of 97.98% and an MCC of 0.962 on an independent validation set, representing improvements of 3.91% in accuracy and 9.60% in MCC over the previous best method, CRISPRCasStack. Conclusions PrePssmCas outperforms seven existing methods (HMMCAS, CASPredict, CRISPRone, CRISPRCasFinder, CRISPRCasTyper, CRISPRloci, and CRISPRCasStack) on an independent validation set. This research contributes valuable computational tools for the classification of CRISPR-Cas systems and provides a promising feature-fusion framework that may support future Cas protein discovery, although its generalization to large-scale and remotely homologous data remains to be established.
Authors
- Qianqian Shi (ORCID: https://orcid.org/0000-0002-5266-6731)
- Chenhan Luo
- Ningyi Zhang
- Yufeng Zhao
- Zhiwei Peng
Publication Details
- Journal
- BMC Bioinformatics
- Published
- 2026-09-25
- DOI
- https://doi.org/10.1186/s12859-026-06569-y
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- article
- Field-Weighted Citation Impact
- 0.00