Tomdok: a fully reproducible, multi-engine virtual screening pipeline with data-driven safety annotation for genome-scale drug repurposing

Virtual screening is a mature but fragmented discipline: a typical campaign requires separate tools for receptor preparation, library curation, protonation, 3D generation, multi-engine docking, scoring, ADMET profiling and figure generation, each with its own formats and undocumented defaults. Most published screens therefore cannot be re-executed from a single command, and safety considerations (hERG liability, PAINS, withdrawn status) are routinely deferred to late stages. This preprint presents Tomdok, an open-source, one-command pipeline (stages 00-09) that integrates target preparation (PDBFixer, pH 7.4, dual receptor formats), ChEMBL 37 library construction with tunable clinical-phase depth, parent-based deduplication and data-driven hERG liability (KCNH2/CHEMBL240 IC50/Ki < 10 uM), multi-GPU GNINA primary screening (exh 4, CNN pose scoring), three-engine consensus refinement (GNINA exh 32, AutoDock Vina exh 32, LeDock; AutoDock 4 evaluated and excluded due to parameter-file incompatibility) with a scale-invariant count-based consensus metric, redocking validation (best/top-pose RMSD), rule-based and data-driven ADMET, CLEAN/FLAGGED safety-aware prioritisation, automated publication figures (PyMOL 3D views and PLIP 2D interaction cards), and generation of a journal-ready supplementary bundle (S0-S9 tables, figures, universal 3D poses, machine-readable parameters, full execution log with rotation). As a case study on the WRN helicase (PDB 7GQU), a synthetic-lethal target in microsatellite-unstable cancers, 3,326 of 3,369 deduplicated approved and Phase-III compounds were docked in 3 h 56 min on two consumer GPUs; the funnel retained 66 primary hits and 60 consensus hits (43 CLEAN / 17 FLAGGED), and self-docking reproduced the crystallographic pose (best-pose RMSD 1.46 A, PASS at the 2.0 A threshold). Two independent full recomputations showed stable top-ranked hits with bounded marginal churn, quantifying GPU floating-point non-determinism in practice. Tomdok provides end-to-end reproducibility (fixed seed 42, resume semantics, log rotation, one-command re-execution for any target) and safety-aware prioritisation for genome-scale repurposing, with code released under MIT and data under CC BY 4.0. Companion biology-first preprint (WRN repurposing candidates): companion deposit, see Related works. Screening data deposit (consensus tables, ADMET profiles, 3D poses, figures, full log): Zenodo, doi:10.5281/zenodo.22109731. Code: Tomdok, https://github.com/NikTomSik/tomdok (MIT); code archive: Zenodo, doi:10.5281/zenodo.22104868.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-08-28
DOI
https://doi.org/10.5281/zenodo.22148096
Primary Topic
Computational Drug Discovery Methods
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Tomdok: a fully reproducible, multi-engine virtual screening pipeline with data-driven safety annotation for genome-scale drug repurposing

Tomas Nikosin
Zenodo (CERN European Organization for Nuclear Research)
Computational Drug Discovery Methods
preprint

Tomdok: a fully reproducible, multi-engine virtual screening pipeline with data-driven safety annotation for genome-scale drug repurposing

Tomas Nikosin
preprint en

Abstract

Virtual screening is a mature but fragmented discipline: a typical campaign requires separate tools for receptor preparation, library curation, protonation, 3D generation, multi-engine docking, scoring, ADMET profiling and figure generation, each with its own formats and undocumented defaults. Most published screens therefore cannot be re-executed from a single command, and safety considerations (hERG liability, PAINS, withdrawn status) are routinely deferred to late stages. This preprint presents Tomdok, an open-source, one-command pipeline (stages 00-09) that integrates target preparation (PDBFixer, pH 7.4, dual receptor formats), ChEMBL 37 library construction with tunable clinical-phase depth, parent-based deduplication and data-driven hERG liability (KCNH2/CHEMBL240 IC50/Ki < 10 uM), multi-GPU GNINA primary screening (exh 4, CNN pose scoring), three-engine consensus refinement (GNINA exh 32, AutoDock Vina exh 32, LeDock; AutoDock 4 evaluated and excluded due to parameter-file incompatibility) with a scale-invariant count-based consensus metric, redocking validation (best/top-pose RMSD), rule-based and data-driven ADMET, CLEAN/FLAGGED safety-aware prioritisation, automated publication figures (PyMOL 3D views and PLIP 2D interaction cards), and generation of a journal-ready supplementary bundle (S0-S9 tables, figures, universal 3D poses, machine-readable parameters, full execution log with rotation). As a case study on the WRN helicase (PDB 7GQU), a synthetic-lethal target in microsatellite-unstable cancers, 3,326 of 3,369 deduplicated approved and Phase-III compounds were docked in 3 h 56 min on two consumer GPUs; the funnel retained 66 primary hits and 60 consensus hits (43 CLEAN / 17 FLAGGED), and self-docking reproduced the crystallographic pose (best-pose RMSD 1.46 A, PASS at the 2.0 A threshold). Two independent full recomputations showed stable top-ranked hits with bounded marginal churn, quantifying GPU floating-point non-determinism in practice. Tomdok provides end-to-end reproducibility (fixed seed 42, resume semantics, log rotation, one-command re-execution for any target) and safety-aware prioritisation for genome-scale repurposing, with code released under MIT and data under CC BY 4.0. Companion biology-first preprint (WRN repurposing candidates): companion deposit, see Related works. Screening data deposit (consensus tables, ADMET profiles, 3D poses, figures, full log): Zenodo, doi:10.5281/zenodo.22109731. Code: Tomdok, https://github.com/NikTomSik/tomdok (MIT); code archive: Zenodo, doi:10.5281/zenodo.22104868.

Zenodo (CERN European Organization for Nuclear Research)
Good health and well-being
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.