Public Database Leakage Distorts Model Rankings in Real-Spectra NMR Structure Elucidation
Abstract Sequence models that translate NMR spectra into molecular structures report up to 96% top-1 accuracy, but they are pretrained on simulated spectra and tested on public databases that can overlap those corpora. We build a leakage-controlled benchmark on nmrshiftdb2 and find that 26.9% of public molecules are exact training-corpus matches under a stated identity rule, with substantial training-set proximity remaining after exact-match removal. On the wider exact-match-removed real-spectrum cohort, released-model top-1 accuracy is 24.4% (95% CI 23.3–25.6%). A training-free rerank using molecular formula and 13C-peak count consistency raises it to 31.1% (29.9–32.4%), supporting leakage-controlled evaluation and decode-time consistency checks before claims of experimental generalization for simulation-pretrained inverse models.
Authors
- Yutao Guo (ORCID: https://orcid.org/0000-0002-7626-7751)
- Dan Wu (ORCID: https://orcid.org/0000-0002-5316-3706)
- Xuezhou Zhao (ORCID: https://orcid.org/0000-0002-9433-8015)
- mengxi Chen
- Zihan Zhang (ORCID: https://orcid.org/0009-0001-6239-8806)
Institutions
- Soochow University (TW)
- St. Stephen's University (CA)
Publication Details
- Journal
- Journal of Chemical Information and Modeling
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1021/acs.jcim.6c02552
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00