PhageScout: Protease Cleavage Site Prediction Using an Experimental Substrate Phage Display Motif-Based Approach
Identification of protease cleavage sites is essential for understanding biological regulation and disease mechanisms, yet many predictive approaches rely on annotated substrates and curated databases, limiting performance for poorly characterized proteases. We present PhageScout, a framework for database-independent generation of protease-specific features to predict cleavage sites using de novo experimental substrate phage display screening. We screened a randomized 5-mer phage display library against two neutrophil serine proteases (cathepsin G, elastase). Cleaved peptides generated position weight matrices (PWMs) and peptide enrichment scores to evaluate cleavage-site likelihood across substrate sequences. Sequence-derived scores were integrated with structural features, including accessibility and flexibility, using XGBoost classification models. Performance was benchmarked against annotated cleavage sites from the MEROPS peptidase database as reference data. Phage-derived PWM scores alone captured protease preferences and discriminated cleavage sites from background sites. Without model fitting, PWM scores achieved an area under the curve (AUC) of 0.756 (95%CI: 0.714–0.797) (cathepsin G) and 0.787 (95%CI: 0.753–0.821) (elastase). Combining broad and specific phage-derived scores improved cathepsin G prediction (AUC = 0.783), whereas this improvement was not observed for elastase. Compared to only phage-derived features, XGBoost models integrating phage sequence and structural features provided modest gains for elastase (AUC = 0.775 to 0.806), with phage-derived features ranking among the strongest predictors, but not cathepsin G (AUC = 0.702 to 0.710). Our findings demonstrate that PhageScout can use experimentally derived cleavage signatures to generate protease-specific predictive features and prioritize protease cleavage sites, providing a framework that warrants further validation across diverse proteases and biological contexts.
Authors
- Eddy Yu
- Colin A. Kretz (ORCID: https://orcid.org/0000-0001-7979-7603)
- Matthew L. Holding (ORCID: https://orcid.org/0000-0003-3477-3012)
- Rex Huang
- Andrew Chan
- Cherie Teney
Institutions
- University of Michigan (US)
- Thrombosis and Atherosclerosis Research Institute (CA)
- Hamilton Health Sciences (CA)
- McMaster University (CA)
Publication Details
- Journal
- International Journal of Molecular Sciences
- Published
- 2026-08-25
- DOI
- https://doi.org/10.3390/ijms27177593
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Institutes of Health
- Canadian Institutes of Health Research
- Natural Sciences and Engineering Research Council of Canada