Benchmarking Experiments with External Data, Domain-Specific Features, and Learning Strategies for Modeling ADME and Potency of Coronavirus Main Protease Inhibitors

Abstract Predictive computational models of potency and pharmacokinetic properties of antiviral drug candidates can dramatically accelerate the path from molecular design to clinical evaluation. The ASAP-Polaris-OpenADMET consortium hosted the open-science Antiviral Drug Discovery Challenge 2025, providing a rigorous setting for this problem and requiring prediction of potency against the SARS-CoV-2 and MERS-CoV Mpro alongside five ADME end points (LogD, KSOL, HLM, MLM, and MDR1-MDCKII). In response, we pursued a domain-aware modeling strategy, reasoning that the distinctive chemical properties of antiviral compounds could be explicitly engineered in the feature space. We tested this premise by isolating the key molecular descriptors capable of identifying antivirals, augmenting challenge data sets with curated public data, and building both classical machine learning and graph neural network (GNN) variants. For ADME modeling, ensemble gradient boosting models trained on our bespoke feature space consistently outperformed both broader tree-based approaches and GNN variants across the five ADME end points in blind evaluation. By a factorial design of experiments and statistical benchmarking, we evaluated the performance of the optimal XGBoost model for each end point. For modeling potency, we explored varied pretraining paradigms and ensemble fusion strategies, and found that a stacking ensemble combining docking-enhanced XGBoost with a stereochemistry-aware AttentiveFP Graph-Transformer model yielded the strongest significant generalization on the Challenge test set, suggesting that orthogonal model families reduced epistemic uncertainty attributable to model selection in data-limited, multitarget settings. Beyond benchmarking, our experiments provided a testbed for evaluating critical assumptions in the field, encompassing data augmentation strategies, splitting protocols, and ensemble integration approaches. Our analysis revealed that a consistently optimal model is elusive, and that several standard industry practices fail to withstand close inspection. Overall, our work provides both practical benchmarked modeling strategies for ADME and potency prediction and a reproducible computational foundation for AI-driven antiviral drug discovery efforts.

Authors

Institutions

Publication Details

Journal
Journal of Chemical Information and Modeling
Published
2026-09-29
DOI
https://doi.org/10.1021/acs.jcim.6c00304
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmarking Experiments with External Data, Domain-Specific Features, and Learning Strategies for Modeling ADME and Potency of Coronavirus Main Protease Inhibitors

Ashok Palaniappan, Ida Titus
Journal of Chemical Information and Modeling
Computational Drug Discovery Methods
article

Benchmarking Experiments with External Data, Domain-Specific Features, and Learning Strategies for Modeling ADME and Potency of Coronavirus Main Protease Inhibitors

Ashok Palaniappan, Ida Titus
article en

Abstract

Abstract Predictive computational models of potency and pharmacokinetic properties of antiviral drug candidates can dramatically accelerate the path from molecular design to clinical evaluation. The ASAP-Polaris-OpenADMET consortium hosted the open-science Antiviral Drug Discovery Challenge 2025, providing a rigorous setting for this problem and requiring prediction of potency against the SARS-CoV-2 and MERS-CoV Mpro alongside five ADME end points (LogD, KSOL, HLM, MLM, and MDR1-MDCKII). In response, we pursued a domain-aware modeling strategy, reasoning that the distinctive chemical properties of antiviral compounds could be explicitly engineered in the feature space. We tested this premise by isolating the key molecular descriptors capable of identifying antivirals, augmenting challenge data sets with curated public data, and building both classical machine learning and graph neural network (GNN) variants. For ADME modeling, ensemble gradient boosting models trained on our bespoke feature space consistently outperformed both broader tree-based approaches and GNN variants across the five ADME end points in blind evaluation. By a factorial design of experiments and statistical benchmarking, we evaluated the performance of the optimal XGBoost model for each end point. For modeling potency, we explored varied pretraining paradigms and ensemble fusion strategies, and found that a stacking ensemble combining docking-enhanced XGBoost with a stereochemistry-aware AttentiveFP Graph-Transformer model yielded the strongest significant generalization on the Challenge test set, suggesting that orthogonal model families reduced epistemic uncertainty attributable to model selection in data-limited, multitarget settings. Beyond benchmarking, our experiments provided a testbed for evaluating critical assumptions in the field, encompassing data augmentation strategies, splitting protocols, and ensemble integration approaches. Our analysis revealed that a consistently optimal model is elusive, and that several standard industry practices fail to withstand close inspection. Overall, our work provides both practical benchmarked modeling strategies for ADME and potency prediction and a reproducible computational foundation for AI-driven antiviral drug discovery efforts.

Journal of Chemical Information and Modeling
SASTRA University (IN)
Industry, innovation and infrastructure
Openalex Percentile: Top 9%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.