CCSBase2: A Unified Benchmark Dataset and Machine Learning Framework for Robust CCS Prediction
Abstract The identification of unknown analytes remains a persistent challenge in untargeted mass spectrometry-based analyses. Ion mobility–mass spectrometry (IM-MS) has expanded identification capabilities by enabling the use of collision cross section (CCS) as an orthogonal molecular identifier, but experimentally measured CCS values remain sparse relative to the breadth of known chemical space. Machine learning (ML) approaches have emerged as a complementary strategy for CCS prediction. Many existing models, however, suffer from poor generalizability and accuracy across structurally diverse chemical classes. Here, we present CCSBase2, a large, unified CCS database comprising 66,153 measurements assembled by integrating multiple public datasets, along with a classical ML framework for CCS prediction using Morgan count fingerprints. CCSBase2 achieved a mean relative error of 1.73%, median relative error of 1.23%, and root mean squared error of 4.89 Å2 on a held-out test set. CCSBase2 was trained on a consumer-grade Apple Silicon M4 Pro CPU, demonstrating that accurate CCS prediction can be achieved with minimal computational resources while maintaining fair performance when tested against out-of-distribution compounds. The final model is available at https://ccsbase.net/ccsbase2, and the codebase and datasets are freely available on GitHub (https://github.com/libinxulab/ccsbase2).
Authors
- Ryan Nguyen (ORCID: https://orcid.org/0000-0003-0953-7536)
- Libin Xu (ORCID: https://orcid.org/0000-0003-1021-5200)
- Griffin Rangel
- Amogh Bantwal
- Reuben Santoso
Institutions
- University of Washington (US)
Publication Details
- Journal
- Analytical Chemistry
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1021/acs.analchem.6c03711
- Primary Topic
- Mass Spectrometry Techniques and Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00