Extracting Reproducibility Metadata from Machine-Learning Papers: a Taxonomy of Silent Errors and the Cost of Human Verification

Reproducing a machine-learning result requires knowing how the experiment was run: the data, the model and its size, the optimiser, the learning rate, the batch size, the number of steps, the hardware and the seeds. This information is scattered through a paper---often only in a table caption or an appendix---and is increasingly extracted automatically by language-model pipelines, whose evaluations report how often they are right. We ask the two questions those evaluations leave open: when a pipeline is wrong, does it know it is wrong; and how much human effort does it take to find out? We build a benchmark of 80 arXiv machine-learning papers and annotate 24 of them (240 field decisions), 12 of them through two independent annotation passes, carrying verbatim, line-numbered evidence for every reported and every absent field; the passes agree on 97.5% of presence/absence decisions (Cohen's κ = 0.94), and the residual disagreements are adjudicated under a written scope rule. On that gold standard we run six extraction conditions of increasing agency---a single untruncated prompt, a tool-using agent, schema constraints, table/appendix re-reading, per-field self-assessment with a suspect flag, and mandatory evidence back-linking---over three difficulty strata, 432 full-paper extractions in total. Four results. First, agency buys no accuracy: present recall varies by one percentage point across all six conditions (0.883–0.894) while wall-clock time varies threefold. Second, honesty is a different quantity: 92% of the 387 wrong values were returned without any indication of doubt, and the share of returned values that are silently wrong falls from 12.7% to 5.5% only in the condition asked to flag its own doubts; requiring an evidence quote for every value restores the silent rate to 11.2% because it crowds out the doubt flag. Third, the errors are not inventions: every returned value that carries a checkable anchor occurs somewhere in the paper, and 375 of the 387 errors are values the document does contain—misattributions that belong to another configuration, another table, or the reference list, or values lifted from elsewhere in the paper—which is why a gold-standard-free anchor audit detects none of them (we report that as a negative result, with the five PDF text-layer artefacts it originally fired on pinned as a 22-case regression test). Fourth, human verification is the binding cost and it is predictable: verifying one field by hand requires a median of 12 candidate passages (up to 273; 4,707 in total for the gold sample) against one passage with an evidence link, although 4 of 25 spot-checked links failed adjudication, three of them because the quoted line does not hold the value. All corpus, prompts, gold annotations, adjudication records, raw runs, per-item API costs and analysis scripts are released with the paper, together with a per-session cost ledger (runs/session-costs.json: the grid itself cost $2.84, the 644 sessions recorded in the released artefacts cost $3.22).

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23118537
Primary Topic
Scientific Computing and Data Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Extracting Reproducibility Metadata from Machine-Learning Papers: a Taxonomy of Silent Errors and the Cost of Human Verification

Heng Li
Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
preprint

Extracting Reproducibility Metadata from Machine-Learning Papers: a Taxonomy of Silent Errors and the Cost of Human Verification

Heng Li
preprint en

Abstract

Reproducing a machine-learning result requires knowing how the experiment was run: the data, the model and its size, the optimiser, the learning rate, the batch size, the number of steps, the hardware and the seeds. This information is scattered through a paper---often only in a table caption or an appendix---and is increasingly extracted automatically by language-model pipelines, whose evaluations report how often they are right. We ask the two questions those evaluations leave open: when a pipeline is wrong, does it know it is wrong; and how much human effort does it take to find out? We build a benchmark of 80 arXiv machine-learning papers and annotate 24 of them (240 field decisions), 12 of them through two independent annotation passes, carrying verbatim, line-numbered evidence for every reported and every absent field; the passes agree on 97.5% of presence/absence decisions (Cohen's κ = 0.94), and the residual disagreements are adjudicated under a written scope rule. On that gold standard we run six extraction conditions of increasing agency---a single untruncated prompt, a tool-using agent, schema constraints, table/appendix re-reading, per-field self-assessment with a suspect flag, and mandatory evidence back-linking---over three difficulty strata, 432 full-paper extractions in total. Four results. First, agency buys no accuracy: present recall varies by one percentage point across all six conditions (0.883–0.894) while wall-clock time varies threefold. Second, honesty is a different quantity: 92% of the 387 wrong values were returned without any indication of doubt, and the share of returned values that are silently wrong falls from 12.7% to 5.5% only in the condition asked to flag its own doubts; requiring an evidence quote for every value restores the silent rate to 11.2% because it crowds out the doubt flag. Third, the errors are not inventions: every returned value that carries a checkable anchor occurs somewhere in the paper, and 375 of the 387 errors are values the document does contain—misattributions that belong to another configuration, another table, or the reference list, or values lifted from elsewhere in the paper—which is why a gold-standard-free anchor audit detects none of them (we report that as a negative result, with the five PDF text-layer artefacts it originally fired on pinned as a 22-case regression test). Fourth, human verification is the binding cost and it is predictable: verifying one field by hand requires a median of 12 candidate passages (up to 273; 4,707 in total for the gold sample) against one passage with an evidence link, although 4 of 25 spot-checked links failed adjudication, three of them because the quoted line does not hold the value. All corpus, prompts, gold annotations, adjudication records, raw runs, per-item API costs and analysis scripts are released with the paper, together with a per-session cost ledger (runs/session-costs.json: the grid itself cost $2.84, the 644 sessions recorded in the released artefacts cost $3.22).

Zenodo (CERN European Organization for Nuclear Research)
Central South University (CN)
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.