ToxBench: a leakage-audited multi-task benchmark for predictive toxicology with calibration, uncertainty, and applicability domain analysis

Machine learning models for toxicity prediction are routinely evaluated using random train/test splits, allowing structurally similar compounds to appear in both sets and inflating reported performance metrics. A rigorous benchmark quantifying this overestimation alongside calibration, uncertainty, and applicability domain analyses is needed to guide practitioners in selecting and trusting toxicity prediction models. We present ToxBench, a leakage-audited benchmark comprising three toxicology datasets (Tox21, 7,538 compounds, 12 tasks; ClinTox, 1,379 compounds, 2 tasks; SIDER, 1,350 compounds, 27 tasks) processed through a transparent standardization pipeline with explicit reporting of conflicting-label removals. Four model classes were evaluated: Random Forest, XGBoost, MLP, and Graph Neural Network (GNN), each trained under random and Bemis-Murcko scaffold-based splits across five independent seeds (120 experimental conditions). Analyses included post-hoc probability calibration, ensemble-based uncertainty quantification, nearest-neighbor applicability domain analysis, and scaffold-level error analysis. Scaffold splitting consistently reduced AUROC by 0.057–0.079 points across all model classes on Tox21 (mean drop: 0.070) and by 0.031–0.035 points on SIDER for three of four models, demonstrating systematic performance overestimation under random splitting. ClinTox showed reversed performance ordering due to small dataset size and extreme class imbalance, with high seed-to-seed variance (± 0.085–0.160) confirming results are dominated by sampling noise. Post-hoc calibration reduced ECE by 67–68% on ClinTox, confirming reliable probability estimates are achievable with simple correction. On Tox21 and SIDER, raw Random Forest predictions were already well-calibrated (ECE 0.018 and 0.057 respectively), with post-hoc methods providing no additional benefit. Scaffold splitting consistently worsened calibration across all three datasets. Applicability domain analysis revealed a consistent positive relationship between nearest-neighbor Tanimoto similarity and predictive reliability, with low-similarity compounds showing notably reduced AUROC under scaffold splitting. Ensemble uncertainty estimates showed promising utility, with confidence-based filtering improving AUROC by up to 0.025 points on SIDER. Random splitting overestimates performance by 0.057–0.079 AUROC points on Tox21 and 0.031–0.035 on SIDER. ToxBench provides a transparent, reproducible, leakage-audited benchmark supporting scaffold splitting, multi-seed evaluation, and applicability domain filtering as standard practice in toxicity prediction benchmarking. Rather than assuming scaffold splitting is leakage-free, we audit residual structural similarity directly: scaffold splitting substantially reduces, but does not eliminate, train–test analog leakage (e.g., the fraction of Tox21 test compounds with a nearest-neighbour Tanimoto > 0.6 to training falls from 43.9% under random splitting to 12.1% under scaffold splitting), and we quantify this residual and provide an applicability-domain filter to flag it.

Authors

Institutions

Publication Details

Journal
BMC Bioinformatics
Published
2026-09-05
DOI
https://doi.org/10.1186/s12859-026-06621-x
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

ToxBench: a leakage-audited multi-task benchmark for predictive toxicology with calibration, uncertainty, and applicability domain analysis

Tae‐Sik Park, Sanggyun Yi, Kartic, Yeeun Seo
BMC Bioinformatics
Computational Drug Discovery Methods
article

ToxBench: a leakage-audited multi-task benchmark for predictive toxicology with calibration, uncertainty, and applicability domain analysis

Tae‐Sik Park, Sanggyun Yi, Kartic, Yeeun Seo
article en

Abstract

Machine learning models for toxicity prediction are routinely evaluated using random train/test splits, allowing structurally similar compounds to appear in both sets and inflating reported performance metrics. A rigorous benchmark quantifying this overestimation alongside calibration, uncertainty, and applicability domain analyses is needed to guide practitioners in selecting and trusting toxicity prediction models. We present ToxBench, a leakage-audited benchmark comprising three toxicology datasets (Tox21, 7,538 compounds, 12 tasks; ClinTox, 1,379 compounds, 2 tasks; SIDER, 1,350 compounds, 27 tasks) processed through a transparent standardization pipeline with explicit reporting of conflicting-label removals. Four model classes were evaluated: Random Forest, XGBoost, MLP, and Graph Neural Network (GNN), each trained under random and Bemis-Murcko scaffold-based splits across five independent seeds (120 experimental conditions). Analyses included post-hoc probability calibration, ensemble-based uncertainty quantification, nearest-neighbor applicability domain analysis, and scaffold-level error analysis. Scaffold splitting consistently reduced AUROC by 0.057–0.079 points across all model classes on Tox21 (mean drop: 0.070) and by 0.031–0.035 points on SIDER for three of four models, demonstrating systematic performance overestimation under random splitting. ClinTox showed reversed performance ordering due to small dataset size and extreme class imbalance, with high seed-to-seed variance (± 0.085–0.160) confirming results are dominated by sampling noise. Post-hoc calibration reduced ECE by 67–68% on ClinTox, confirming reliable probability estimates are achievable with simple correction. On Tox21 and SIDER, raw Random Forest predictions were already well-calibrated (ECE 0.018 and 0.057 respectively), with post-hoc methods providing no additional benefit. Scaffold splitting consistently worsened calibration across all three datasets. Applicability domain analysis revealed a consistent positive relationship between nearest-neighbor Tanimoto similarity and predictive reliability, with low-similarity compounds showing notably reduced AUROC under scaffold splitting. Ensemble uncertainty estimates showed promising utility, with confidence-based filtering improving AUROC by up to 0.025 points on SIDER. Random splitting overestimates performance by 0.057–0.079 AUROC points on Tox21 and 0.031–0.035 on SIDER. ToxBench provides a transparent, reproducible, leakage-audited benchmark supporting scaffold splitting, multi-seed evaluation, and applicability domain filtering as standard practice in toxicity prediction benchmarking. Rather than assuming scaffold splitting is leakage-free, we audit residual structural similarity directly: scaffold splitting substantially reduces, but does not eliminate, train–test analog leakage (e.g., the fraction of Tox21 test compounds with a nearest-neighbour Tanimoto > 0.6 to training falls from 43.9% under random splitting to 12.1% under scaffold splitting), and we quantify this residual and provide an applicability-domain filter to flag it.

BMC Bioinformatics
Gachon University (KR)
Gachon University
Openalex Percentile: Top 8%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.