Activity Landscape Roughness Anticipates Machine-Learning Reliability Across Environmental Chemistry Endpoints

Abstract Machine learning increasingly supports chemical risk assessment under REACH, TSCA, and OECD QSAR frameworks, but random-split performance can be optimistic for structurally novel chemicals. We ask whether task difficulty can be anticipated before model training. EnvMolBench spans 45 datasets, 114,500+ end point-specific records (55,984 unique structures), and 6500+ model–dataset combinations from 13 algorithms and 6 representation categories. Training-set activity landscape roughness, quantified by nearest-neighbor disagreement rate (classification) or the Structure–Activity Landscape Index (regression), was associated with best-observed performance (Spearman ρ = −0.71, 95% CI [−0.94, −0.26], p = 0.0009 for classification; ρ = −0.47, 95% CI [−0.81, −0.04], p = 0.019 for regression), surviving family-wide FDR correction for classification (q = 0.0085) but not regression (q = 0.063). On high-roughness end points, excluding activity-cliff compounds post hoc raised AUC by a mean 0.134 and lowered standardized RMSE by 0.067, whereas removing cliffs from training did not help. The three evaluated pretrained models did not outperform well-tuned baselines on the ten roughest end points. Roughness therefore provides an empirical, representation-relative diagnostic of achievable performance rather than a fundamental ceiling, motivating a workflow linking it to method selection, applicability-domain assessment, and conformal calibration. EnvMolBench, its datasets, splits, and baselines are released openly.

Authors

Institutions

Publication Details

Journal
Environmental Science & Technology
Published
2026-09-07
DOI
https://doi.org/10.1021/acs.est.6c06820
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Activity Landscape Roughness Anticipates Machine-Learning Reliability Across Environmental Chemistry Endpoints

Shifa Zhong, Zeting Wu, Haoyang Li
Environmental Science & Technology
Computational Drug Discovery Methods
article

Activity Landscape Roughness Anticipates Machine-Learning Reliability Across Environmental Chemistry Endpoints

Shifa Zhong, Zeting Wu, Haoyang Li
article en

Abstract

Abstract Machine learning increasingly supports chemical risk assessment under REACH, TSCA, and OECD QSAR frameworks, but random-split performance can be optimistic for structurally novel chemicals. We ask whether task difficulty can be anticipated before model training. EnvMolBench spans 45 datasets, 114,500+ end point-specific records (55,984 unique structures), and 6500+ model–dataset combinations from 13 algorithms and 6 representation categories. Training-set activity landscape roughness, quantified by nearest-neighbor disagreement rate (classification) or the Structure–Activity Landscape Index (regression), was associated with best-observed performance (Spearman ρ = −0.71, 95% CI [−0.94, −0.26], p = 0.0009 for classification; ρ = −0.47, 95% CI [−0.81, −0.04], p = 0.019 for regression), surviving family-wide FDR correction for classification (q = 0.0085) but not regression (q = 0.063). On high-roughness end points, excluding activity-cliff compounds post hoc raised AUC by a mean 0.134 and lowered standardized RMSE by 0.067, whereas removing cliffs from training did not help. The three evaluated pretrained models did not outperform well-tuned baselines on the ten roughest end points. Roughness therefore provides an empirical, representation-relative diagnostic of achievable performance rather than a fundamental ceiling, motivating a workflow linking it to method selection, applicability-domain assessment, and conformal calibration. EnvMolBench, its datasets, splits, and baselines are released openly.

Environmental Science & Technology
Tongji University (CN), East China Normal University (CN)
National Natural Science Foundation of China
Openalex Percentile: Top 9%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Activity Landscape Roughness Anticipates Machine-Learning Reliability Across Environmental Chemistry Endpoints — Shifa Zhong, Zeting Wu, et al. · Environmental Science & Technology (2026) | TGRS Research Map | TGRS