Developing SCL2205 : A Protein Sequence-based Spatial Modelling Dataset for the Protein Language Model Frontier

MOTIVATION: Deep learning (DL) has substantially advanced protein subcellular localisation (SCL) prediction, yet its potential remains constrained by suboptimal input preparation and limited high-quality reference data. Furthermore, existing state-of-the-art (SoTA) predictors suffer from performance metric inflation due to unmitigated training-to-testing data leakage during homology augmentation. We address these challenges by introducing SCL2205, a leak-minimised benchmark dataset and pipeline curated specifically to support trustworthy, scalable, and reproducible DL-based SCL modelling. RESULTS: SCL2205 was constructed from the universal protein knowledgebase (UniProtKB) using rigorous preprocessing, manual label mapping, and stringent partitioning. When evaluated on independent test sets, SCL2205 yielded up to a 10.8 percentage point improvement in macro area under the precision-recall curve (PR-AUC) over SoTA baselines (mean Δ 95% CI=0.07-0.12 ), with maximum benefits observed when paired with modern protein language models (PLMs). Crucially, we quantify for the first time a systemic 5.2%±0.32 data leakage rate in conventional homology augmentation workflows-even when restricting sequence similarity searches to just 10% of the training set. AVAILABILITY AND IMPLEMENTATION: The dataset is openly available on Dryad under a CC0 1.0 Universal licence (https://doi.org/10.5061/dryad.2ngf1vj1t). The dataset interface is available as an installable Python package, p-scldata (v2026.2.0), under the MIT licence on the Python Package Index (PyPI). Full code and data repositories are hosted on GitHub (https://github.com/ousodaniel/scldata) and archived on Zenodo (https://doi.org/10.5281/zenodo.21796423). SUPPLEMENTARY INFORMATION: Supplementary File S1 contains code snippets, per-class PR-AUC breakdowns, class prevalence details, statistical comparison tests, and supplementary figures. Supplementary File S2 contains the exact mapping used in curation.

Authors

Institutions

Publication Details

Journal
Bioinformatics
Published
2026-09-17
DOI
https://doi.org/10.1093/bioinformatics/btag666
Primary Topic
Machine Learning in Bioinformatics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Developing SCL2205 : A Protein Sequence-based Spatial Modelling Dataset for the Protein Language Model Frontier

Gianluca Pollastri, Daniel O. Ouso
Bioinformatics
Machine Learning in Bioinformatics
article

Developing SCL2205 : A Protein Sequence-based Spatial Modelling Dataset for the Protein Language Model Frontier

Gianluca Pollastri, Daniel O. Ouso
article en

Abstract

MOTIVATION: Deep learning (DL) has substantially advanced protein subcellular localisation (SCL) prediction, yet its potential remains constrained by suboptimal input preparation and limited high-quality reference data. Furthermore, existing state-of-the-art (SoTA) predictors suffer from performance metric inflation due to unmitigated training-to-testing data leakage during homology augmentation. We address these challenges by introducing SCL2205, a leak-minimised benchmark dataset and pipeline curated specifically to support trustworthy, scalable, and reproducible DL-based SCL modelling. RESULTS: SCL2205 was constructed from the universal protein knowledgebase (UniProtKB) using rigorous preprocessing, manual label mapping, and stringent partitioning. When evaluated on independent test sets, SCL2205 yielded up to a 10.8 percentage point improvement in macro area under the precision-recall curve (PR-AUC) over SoTA baselines (mean Δ 95% CI=0.07-0.12 ), with maximum benefits observed when paired with modern protein language models (PLMs). Crucially, we quantify for the first time a systemic 5.2%±0.32 data leakage rate in conventional homology augmentation workflows-even when restricting sequence similarity searches to just 10% of the training set. AVAILABILITY AND IMPLEMENTATION: The dataset is openly available on Dryad under a CC0 1.0 Universal licence (https://doi.org/10.5061/dryad.2ngf1vj1t). The dataset interface is available as an installable Python package, p-scldata (v2026.2.0), under the MIT licence on the Python Package Index (PyPI). Full code and data repositories are hosted on GitHub (https://github.com/ousodaniel/scldata) and archived on Zenodo (https://doi.org/10.5281/zenodo.21796423). SUPPLEMENTARY INFORMATION: Supplementary File S1 contains code snippets, per-class PR-AUC breakdowns, class prevalence details, statistical comparison tests, and supplementary figures. Supplementary File S2 contains the exact mapping used in curation.

Bioinformatics
University College Dublin (IE), Ollscoil na Gaillimhe – University of Galway (IE)
Openalex Percentile: Top 18%
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.