A unified framework for the systematic detection of non-canonical proteins in standard proteomics workflows using a Ribo-seq informed transcriptomic language model

Abstract Although ribosome profiling (Ribo-seq) has revealed widespread translation of non-canonical open reading frames (ncORFs), integrating this information into routine mass spectrometry (MS) workflows remains challenging due to biochemical constraints and large search databases. We retrained TIS Transformer, a transcriptomic language model, by incorporating 40,397 Ribo-seq-derived non-canonical translation initiation sites, thereby enabling the prediction of 48,265 ncORFs, including non-AUG starts, with only a minor impact on canonical protein detection. We combined these predictions with Swiss-Prot sequences to create Swiss-Prot/ncProt, a size-controlled database suitable for standard proteomic analysis. Reanalysis of three independent MS datasets using this database maintained over 98% concordance with Swiss-Prot searches, identifying 541 non-canonical proteins, 429 of which had orthogonal evidence of their existence. This unified framework facilitates systematic detection of non-canonical proteins in conventional proteomics workflows, bridging the gap between canonical and non-canonical proteome analyses without requiring specialized experimental protocols.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-09-25
DOI
https://doi.org/10.1038/s41598-026-72489-9
Primary Topic
Advanced Proteomics Techniques and Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A unified framework for the systematic detection of non-canonical proteins in standard proteomics workflows using a Ribo-seq informed transcriptomic language model

Xavier Roucou, Sébastien Leblanc, Nicolas Provencher, Jean-François Jacques
Scientific Reports
Advanced Proteomics Techniques and Applications
article

A unified framework for the systematic detection of non-canonical proteins in standard proteomics workflows using a Ribo-seq informed transcriptomic language model

Xavier Roucou, Sébastien Leblanc, Nicolas Provencher, Jean-François Jacques
article en

Abstract

Abstract Although ribosome profiling (Ribo-seq) has revealed widespread translation of non-canonical open reading frames (ncORFs), integrating this information into routine mass spectrometry (MS) workflows remains challenging due to biochemical constraints and large search databases. We retrained TIS Transformer, a transcriptomic language model, by incorporating 40,397 Ribo-seq-derived non-canonical translation initiation sites, thereby enabling the prediction of 48,265 ncORFs, including non-AUG starts, with only a minor impact on canonical protein detection. We combined these predictions with Swiss-Prot sequences to create Swiss-Prot/ncProt, a size-controlled database suitable for standard proteomic analysis. Reanalysis of three independent MS datasets using this database maintained over 98% concordance with Swiss-Prot searches, identifying 541 non-canonical proteins, 429 of which had orthogonal evidence of their existence. This unified framework facilitates systematic detection of non-canonical proteins in conventional proteomics workflows, bridging the gap between canonical and non-canonical proteome analyses without requiring specialized experimental protocols.

Scientific Reports
Quality Education
Openalex Percentile: Top 23%
Advanced Proteomics Techniques and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.