Tomdok: a fully reproducible, multi-engine virtual screening pipeline with data-driven safety annotation for genome-scale drug repurposing
Virtual screening is a mature but fragmented discipline: a typical campaign requires separate tools for receptor preparation, library curation, protonation, 3D generation, multi-engine docking, scoring, ADMET profiling and figure generation, each with its own formats and undocumented defaults. Most published screens therefore cannot be re-executed from a single command, and safety considerations (hERG liability, PAINS, withdrawn status) are routinely deferred to late stages. This preprint presents Tomdok, an open-source, one-command pipeline (stages 00-09) that integrates target preparation (PDBFixer, pH 7.4, dual receptor formats), ChEMBL 37 library construction with tunable clinical-phase depth, parent-based deduplication and data-driven hERG liability (KCNH2/CHEMBL240 IC50/Ki < 10 uM), multi-GPU GNINA primary screening (exh 4, CNN pose scoring), three-engine consensus refinement (GNINA exh 32, AutoDock Vina exh 32, LeDock; AutoDock 4 evaluated and excluded due to parameter-file incompatibility) with a scale-invariant count-based consensus metric, redocking validation (best/top-pose RMSD), rule-based and data-driven ADMET, CLEAN/FLAGGED safety-aware prioritisation, automated publication figures (PyMOL 3D views and PLIP 2D interaction cards), and generation of a journal-ready supplementary bundle (S0-S9 tables, figures, universal 3D poses, machine-readable parameters, full execution log with rotation). As a case study on the WRN helicase (PDB 7GQU), a synthetic-lethal target in microsatellite-unstable cancers, 3,326 of 3,369 deduplicated approved and Phase-III compounds were docked in 3 h 56 min on two consumer GPUs; the funnel retained 66 primary hits and 60 consensus hits (43 CLEAN / 17 FLAGGED), and self-docking reproduced the crystallographic pose (best-pose RMSD 1.46 A, PASS at the 2.0 A threshold). Two independent full recomputations showed stable top-ranked hits with bounded marginal churn, quantifying GPU floating-point non-determinism in practice. Tomdok provides end-to-end reproducibility (fixed seed 42, resume semantics, log rotation, one-command re-execution for any target) and safety-aware prioritisation for genome-scale repurposing, with code released under MIT and data under CC BY 4.0. Companion biology-first preprint (WRN repurposing candidates): companion deposit, see Related works. Screening data deposit (consensus tables, ADMET profiles, 3D poses, figures, full log): Zenodo, doi:10.5281/zenodo.22109731. Code: Tomdok, https://github.com/NikTomSik/tomdok (MIT); code archive: Zenodo, doi:10.5281/zenodo.22104868.
Authors
- Tomas Nikosin (ORCID: https://orcid.org/0009-0003-5804-6589)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-08-28
- DOI
- https://doi.org/10.5281/zenodo.22148096
- Primary Topic
- Computational Drug Discovery Methods
- Type
- preprint