Pre-registration: what does a generative campaign produce when its reward provably cannot rank poses?

We validated the computational tools at the allosteric doorstop pocket of Schistosoma thioredoxin glutathione reductase before designing anything, and both candidate reward functions failed. AutoDock Vina's score does not select a correct pose for any of the nine crystallographically observed ligands (0 of 9 within 2.00 A) and is anti-correlated with pose quality (rho = +0.216, p = 0.0066). A ligand-based shape, pharmacophore and selectivity objective also selects a correct pose for 0 of 9, admits no passing reweighting across all fifteen non-empty subsets of its terms, and separates the nine references from 886 property-matched decoys at an area under the curve of 0.5303 against a 0.6592 significance bar. Supplying the crystallographic bridging waters, using coordinates no prospective protocol could possess, recovers 0 of 6. Almost nobody validates a generative reward before running the campaign. Having done so, we are in the unusual position of being able to ask what such a campaign actually produces. This document pre-registers that experiment. It asks whether a generative model driven by a reward with no pose-ranking ability exhibits reward hacking, and if so in what specific and measurable way. Four pathology signatures are fixed in advance with numeric thresholds, three independent seeds per arm are specified with all three to be reported, and a control arm using a pose-free reward is specified so that any pathology can be attributed to the reward rather than to the generator. The decision rule, the claim ceiling on any molecule produced, and the analyses that are not permitted after seeing output are all stated. Deposited before any molecule was generated.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-11
DOI
https://doi.org/10.5281/zenodo.23288941
Primary Topic
Computational Drug Discovery Methods
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Pre-registration: what does a generative campaign produce when its reward provably cannot rank poses?

Donmig Jaula Cunanan, Dianne Jaula Cunanan, Andrea Gail O. Mollasgo, Aldrin Bien B. Coloma et al.
Zenodo (CERN European Organization for Nuclear Research)
Computational Drug Discovery Methods
preprint

Pre-registration: what does a generative campaign produce when its reward provably cannot rank poses?

Donmig Jaula Cunanan, Dianne Jaula Cunanan, Andrea Gail O. Mollasgo, Aldrin Bien B. Coloma, Jom Lui B. Mendoza, Myke Nathaniel A. Dungo, Jamie Adrienne C. Fellores, Chelsea Graciel H. Andam, Gian Carlo I. Pangilinan, Cyril D. Religioso, Ajiecha B. Ventura, Leila Marie R. Gomez, Joss Nickquiel M. Goma, Genelle Marie G. Aleganza, Matt Brian P. Ong, Brent Y. Sison, Yelka R. Yap, Hans Christian M. Cablayan, Joanna Michelle A. Almeda, Gregory Dominic L. Zenarosa, Thomas Anthony D. Montemayor, Jade Kylie D. Villaflores
preprint en

Abstract

We validated the computational tools at the allosteric doorstop pocket of Schistosoma thioredoxin glutathione reductase before designing anything, and both candidate reward functions failed. AutoDock Vina's score does not select a correct pose for any of the nine crystallographically observed ligands (0 of 9 within 2.00 A) and is anti-correlated with pose quality (rho = +0.216, p = 0.0066). A ligand-based shape, pharmacophore and selectivity objective also selects a correct pose for 0 of 9, admits no passing reweighting across all fifteen non-empty subsets of its terms, and separates the nine references from 886 property-matched decoys at an area under the curve of 0.5303 against a 0.6592 significance bar. Supplying the crystallographic bridging waters, using coordinates no prospective protocol could possess, recovers 0 of 6. Almost nobody validates a generative reward before running the campaign. Having done so, we are in the unusual position of being able to ask what such a campaign actually produces. This document pre-registers that experiment. It asks whether a generative model driven by a reward with no pose-ranking ability exhibits reward hacking, and if so in what specific and measurable way. Four pathology signatures are fixed in advance with numeric thresholds, three independent seeds per arm are specified with all three to be reported, and a control arm using a pose-free reward is specified so that any pathology can be attributed to the reward rather than to the generator. The decision rule, the claim ceiling on any molecule produced, and the analyses that are not permitted after seeing output are all stated. Deposited before any molecule was generated.

Zenodo (CERN European Organization for Nuclear Research)
Mapúa University (PH)
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.