PrePssmCas: fusing pre-trained features and PSSM features for Cas protein classification

Abstract Background Clustered regularly interspaced short palindromic repeats (CRISPR) and their associated (Cas) proteins play a vital role in prokaryotic adaptive immune systems, defending against foreign genetic material. However, the classification of CRISPR-Cas systems remains challenging due to limited direct identification techniques. Distinct classes of Cas proteins perform different roles, offering an opportunity for system classification through their identification. The accurate identification of specific Cas proteins is therefore essential for advancing CRISPR-Cas system classification. Results To enhance Cas protein discrimination, two complementary types of features were extracted from protein sequences: pre-trained protein language model embeddings and fixed-dimensional features derived from position-specific scoring matrix (PSSM) profiles. Through systematic comparison of five pre-trained models (Prot_BERT, ALBERT, ProtXLNet, ProtT5, and ESM1b) and six PSSM variants, combined with an attention-based aggregation strategy, ESM1b features combined with RPM-PSSM features demonstrated the most comprehensive representation of Cas proteins. By employing random forest-based feature selection, a 143-dimensional feature vector was obtained, comprising 87 pre-trained features and 56 RPM-PSSM features. The selected features achieved an accuracy of 97.98% and an MCC of 0.962 on an independent validation set, representing improvements of 3.91% in accuracy and 9.60% in MCC over the previous best method, CRISPRCasStack. Conclusions PrePssmCas outperforms seven existing methods (HMMCAS, CASPredict, CRISPRone, CRISPRCasFinder, CRISPRCasTyper, CRISPRloci, and CRISPRCasStack) on an independent validation set. This research contributes valuable computational tools for the classification of CRISPR-Cas systems and provides a promising feature-fusion framework that may support future Cas protein discovery, although its generalization to large-scale and remotely homologous data remains to be established.

Authors

Publication Details

Journal
BMC Bioinformatics
Published
2026-09-25
DOI
https://doi.org/10.1186/s12859-026-06569-y
Primary Topic
Machine Learning in Bioinformatics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

PrePssmCas: fusing pre-trained features and PSSM features for Cas protein classification

Qianqian Shi, Chenhan Luo, Ningyi Zhang, Yufeng Zhao et al.
BMC Bioinformatics
Machine Learning in Bioinformatics
article

PrePssmCas: fusing pre-trained features and PSSM features for Cas protein classification

Qianqian Shi, Chenhan Luo, Ningyi Zhang, Yufeng Zhao, Zhiwei Peng
article en

Abstract

Abstract Background Clustered regularly interspaced short palindromic repeats (CRISPR) and their associated (Cas) proteins play a vital role in prokaryotic adaptive immune systems, defending against foreign genetic material. However, the classification of CRISPR-Cas systems remains challenging due to limited direct identification techniques. Distinct classes of Cas proteins perform different roles, offering an opportunity for system classification through their identification. The accurate identification of specific Cas proteins is therefore essential for advancing CRISPR-Cas system classification. Results To enhance Cas protein discrimination, two complementary types of features were extracted from protein sequences: pre-trained protein language model embeddings and fixed-dimensional features derived from position-specific scoring matrix (PSSM) profiles. Through systematic comparison of five pre-trained models (Prot_BERT, ALBERT, ProtXLNet, ProtT5, and ESM1b) and six PSSM variants, combined with an attention-based aggregation strategy, ESM1b features combined with RPM-PSSM features demonstrated the most comprehensive representation of Cas proteins. By employing random forest-based feature selection, a 143-dimensional feature vector was obtained, comprising 87 pre-trained features and 56 RPM-PSSM features. The selected features achieved an accuracy of 97.98% and an MCC of 0.962 on an independent validation set, representing improvements of 3.91% in accuracy and 9.60% in MCC over the previous best method, CRISPRCasStack. Conclusions PrePssmCas outperforms seven existing methods (HMMCAS, CASPredict, CRISPRone, CRISPRCasFinder, CRISPRCasTyper, CRISPRloci, and CRISPRCasStack) on an independent validation set. This research contributes valuable computational tools for the classification of CRISPR-Cas systems and provides a promising feature-fusion framework that may support future Cas protein discovery, although its generalization to large-scale and remotely homologous data remains to be established.

BMC Bioinformatics
Reduced inequalities, Peace, Justice and strong institutions
Openalex Percentile: Top 19%
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

PrePssmCas: fusing pre-trained features and PSSM features for Cas protein classification — Qianqian Shi, Chenhan Luo, et al. · BMC Bioinformatics (2026) | TGRS Research Map | TGRS