Developing SCL2205 : A Protein Sequence-based Spatial Modelling Dataset for the Protein Language Model Frontier
MOTIVATION: Deep learning (DL) has substantially advanced protein subcellular localisation (SCL) prediction, yet its potential remains constrained by suboptimal input preparation and limited high-quality reference data. Furthermore, existing state-of-the-art (SoTA) predictors suffer from performance metric inflation due to unmitigated training-to-testing data leakage during homology augmentation. We address these challenges by introducing SCL2205, a leak-minimised benchmark dataset and pipeline curated specifically to support trustworthy, scalable, and reproducible DL-based SCL modelling. RESULTS: SCL2205 was constructed from the universal protein knowledgebase (UniProtKB) using rigorous preprocessing, manual label mapping, and stringent partitioning. When evaluated on independent test sets, SCL2205 yielded up to a 10.8 percentage point improvement in macro area under the precision-recall curve (PR-AUC) over SoTA baselines (mean Δ 95% CI=0.07-0.12 ), with maximum benefits observed when paired with modern protein language models (PLMs). Crucially, we quantify for the first time a systemic 5.2%±0.32 data leakage rate in conventional homology augmentation workflows-even when restricting sequence similarity searches to just 10% of the training set. AVAILABILITY AND IMPLEMENTATION: The dataset is openly available on Dryad under a CC0 1.0 Universal licence (https://doi.org/10.5061/dryad.2ngf1vj1t). The dataset interface is available as an installable Python package, p-scldata (v2026.2.0), under the MIT licence on the Python Package Index (PyPI). Full code and data repositories are hosted on GitHub (https://github.com/ousodaniel/scldata) and archived on Zenodo (https://doi.org/10.5281/zenodo.21796423). SUPPLEMENTARY INFORMATION: Supplementary File S1 contains code snippets, per-class PR-AUC breakdowns, class prevalence details, statistical comparison tests, and supplementary figures. Supplementary File S2 contains the exact mapping used in curation.
Authors
- Gianluca Pollastri (ORCID: https://orcid.org/0000-0002-5825-4949)
- Daniel O. Ouso (ORCID: https://orcid.org/0000-0003-0994-2558)
Institutions
- University College Dublin (IE)
- Ollscoil na Gaillimhe – University of Galway (IE)
Publication Details
- Journal
- Bioinformatics
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1093/bioinformatics/btag666
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- article
- Field-Weighted Citation Impact
- 0.00