A unified framework for the systematic detection of non-canonical proteins in standard proteomics workflows using a Ribo-seq informed transcriptomic language model
Abstract Although ribosome profiling (Ribo-seq) has revealed widespread translation of non-canonical open reading frames (ncORFs), integrating this information into routine mass spectrometry (MS) workflows remains challenging due to biochemical constraints and large search databases. We retrained TIS Transformer, a transcriptomic language model, by incorporating 40,397 Ribo-seq-derived non-canonical translation initiation sites, thereby enabling the prediction of 48,265 ncORFs, including non-AUG starts, with only a minor impact on canonical protein detection. We combined these predictions with Swiss-Prot sequences to create Swiss-Prot/ncProt, a size-controlled database suitable for standard proteomic analysis. Reanalysis of three independent MS datasets using this database maintained over 98% concordance with Swiss-Prot searches, identifying 541 non-canonical proteins, 429 of which had orthogonal evidence of their existence. This unified framework facilitates systematic detection of non-canonical proteins in conventional proteomics workflows, bridging the gap between canonical and non-canonical proteome analyses without requiring specialized experimental protocols.
Authors
- Xavier Roucou (ORCID: https://orcid.org/0000-0001-9370-5584)
- Sébastien Leblanc (ORCID: https://orcid.org/0000-0003-2599-6716)
- Nicolas Provencher
- Jean-François Jacques
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-25
- DOI
- https://doi.org/10.1038/s41598-026-72489-9
- Primary Topic
- Advanced Proteomics Techniques and Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00