Extracting Reproducibility Metadata from Machine-Learning Papers: a Taxonomy of Silent Errors and the Cost of Human Verification
Reproducing a machine-learning result requires knowing how the experiment was run: the data, the model and its size, the optimiser, the learning rate, the batch size, the number of steps, the hardware and the seeds. This information is scattered through a paper---often only in a table caption or an appendix---and is increasingly extracted automatically by language-model pipelines, whose evaluations report how often they are right. We ask the two questions those evaluations leave open: when a pipeline is wrong, does it know it is wrong; and how much human effort does it take to find out? We build a benchmark of 80 arXiv machine-learning papers and annotate 24 of them (240 field decisions), 12 of them through two independent annotation passes, carrying verbatim, line-numbered evidence for every reported and every absent field; the passes agree on 97.5% of presence/absence decisions (Cohen's κ = 0.94), and the residual disagreements are adjudicated under a written scope rule. On that gold standard we run six extraction conditions of increasing agency---a single untruncated prompt, a tool-using agent, schema constraints, table/appendix re-reading, per-field self-assessment with a suspect flag, and mandatory evidence back-linking---over three difficulty strata, 432 full-paper extractions in total. Four results. First, agency buys no accuracy: present recall varies by one percentage point across all six conditions (0.883–0.894) while wall-clock time varies threefold. Second, honesty is a different quantity: 92% of the 387 wrong values were returned without any indication of doubt, and the share of returned values that are silently wrong falls from 12.7% to 5.5% only in the condition asked to flag its own doubts; requiring an evidence quote for every value restores the silent rate to 11.2% because it crowds out the doubt flag. Third, the errors are not inventions: every returned value that carries a checkable anchor occurs somewhere in the paper, and 375 of the 387 errors are values the document does contain—misattributions that belong to another configuration, another table, or the reference list, or values lifted from elsewhere in the paper—which is why a gold-standard-free anchor audit detects none of them (we report that as a negative result, with the five PDF text-layer artefacts it originally fired on pinned as a 22-case regression test). Fourth, human verification is the binding cost and it is predictable: verifying one field by hand requires a median of 12 candidate passages (up to 273; 4,707 in total for the gold sample) against one passage with an evidence link, although 4 of 25 spot-checked links failed adjudication, three of them because the quoted line does not hold the value. All corpus, prompts, gold annotations, adjudication records, raw runs, per-item API costs and analysis scripts are released with the paper, together with a per-session cost ledger (runs/session-costs.json: the grid itself cost $2.84, the 644 sessions recorded in the released artefacts cost $3.22).
Authors
- Heng Li (ORCID: https://orcid.org/0000-0003-4874-2874)
Institutions
- Central South University (CN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23118538
- Primary Topic
- Scientific Computing and Data Management
- Type
- preprint