Docking-score landscapes shape active-learning performance across Vina, Glide, and SILCS

The rapid expansion of large chemical libraries has created a need for virtual screening workflows that are both efficient and accurate. Active learning (AL) offers a scalable strategy by iteratively training surrogate models to prioritize promising compounds and reduce the number of required docking calculations. However, direct benchmarking of active-learning protocols across docking engines remains limited. In this study, we ask whether docking scores generated by different docking engines affect AL performance and what factors underlie these differences. To do so, we compare four active-learning virtual screening workflows, Vina-MolPAL, Glide-MolPAL, SILCS-MolPAL, and Schrödinger's active-learning Glide, across multiple protein targets and library sizes. Performance was assessed by recovery of top-scoring molecules according to each workflow's respective docking engine, score-prediction accuracy, chemical diversity, and computational cost. Vina-MolPAL achieved the highest top-1% recovery at a 1% batch size, whereas SILCS-MolPAL achieved comparable recovery at a larger batch size. Latent-embedding analyses suggest that the docking-score landscape strongly influences active-learning performance. In addition, combining active learning with SILCS provides a computationally efficient, membrane-aware approach for screening compounds at transmembrane binding sites.

Authors

Institutions

Publication Details

Journal
Journal of Computer-Aided Molecular Design
Published
2026-09-11
DOI
https://doi.org/10.1007/s10822-026-00927-x
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Docking-score landscapes shape active-learning performance across Vina, Glide, and SILCS

Sai Kosaraju, Sunhwan Jo, Alexander D. MacKerell, Yi‐Chun Lin et al.
Journal of Computer-Aided Molecular Design
Computational Drug Discovery Methods
article

Docking-score landscapes shape active-learning performance across Vina, Glide, and SILCS

Sai Kosaraju, Sunhwan Jo, Alexander D. MacKerell, Yi‐Chun Lin, Aashish Bhatt, Yun Luo, Jacob Ede Levine, Joseph Chung, Mingtian Zhao
article en

Abstract

The rapid expansion of large chemical libraries has created a need for virtual screening workflows that are both efficient and accurate. Active learning (AL) offers a scalable strategy by iteratively training surrogate models to prioritize promising compounds and reduce the number of required docking calculations. However, direct benchmarking of active-learning protocols across docking engines remains limited. In this study, we ask whether docking scores generated by different docking engines affect AL performance and what factors underlie these differences. To do so, we compare four active-learning virtual screening workflows, Vina-MolPAL, Glide-MolPAL, SILCS-MolPAL, and Schrödinger's active-learning Glide, across multiple protein targets and library sizes. Performance was assessed by recovery of top-scoring molecules according to each workflow's respective docking engine, score-prediction accuracy, chemical diversity, and computational cost. Vina-MolPAL achieved the highest top-1% recovery at a 1% batch size, whereas SILCS-MolPAL achieved comparable recovery at a larger batch size. Latent-embedding analyses suggest that the docking-score landscape strongly influences active-learning performance. In addition, combining active learning with SILCS provides a computationally efficient, membrane-aware approach for screening compounds at transmembrane binding sites.

Journal of Computer-Aided Molecular DesignVol. 40(1)
University of Maryland, Baltimore (US), Western University of Health Sciences (US), California State Polytechnic University (US)
Western University of Health Sciences
Openalex Percentile: Top 9%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.