Who Grades the Grader? The Verification Gap in Self-Testing Code Agents

When a coding agent writes both code and tests, the two can agree while sharing the same mistake. Passing those tests may then conceal a misunderstood requirement. We measure how often code reported as successful fails tests that were not provided to the agent. We call this mismatch the verification gap. We evaluate one test-first pipeline on selected HumanEval and LiveCodeBench problems and on RGRBench, a collection of 89 Python packages. With one inexpensive model writing tests, generating code, and reviewing both, the withheld tests reject 18.5% of claimed successes on HumanEval, 28.9% on LiveCodeBench, and 41.5% on RGRBench. We then test stronger models as test authors and reviewers while keeping the code generator fixed. Across three providers on the single-function benchmarks, some settings produce fewer success claims that fail the withheld tests, but the evidence for a reduction remains inconclusive. In our main LiveCodeBench experiment, a stronger reviewer also helps the pipeline produce more solutions that pass the withheld tests. On RGRBench, the same reviewer stops many tasks before code is written and produces fewer passing solutions. Failures remain among claimed successes in every configuration studied. The practical lesson is to measure both what an agent produces and how reliably it judges its own work. That judgment also depends on a clear task: some benchmark failures reflect requirements that omit details demanded by the tests. We release the pipeline, recorded runs, and analysis tools to support independent checks of coding agents' success claims.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23226464
Primary Topic
Software Engineering Research
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Who Grades the Grader? The Verification Gap in Self-Testing Code Agents

Brian Sam-Bodden, Eric Kasper, Josh Meyers, Jyoti Vasudev
Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
preprint

Who Grades the Grader? The Verification Gap in Self-Testing Code Agents

Brian Sam-Bodden, Eric Kasper, Josh Meyers, Jyoti Vasudev
preprint en

Abstract

When a coding agent writes both code and tests, the two can agree while sharing the same mistake. Passing those tests may then conceal a misunderstood requirement. We measure how often code reported as successful fails tests that were not provided to the agent. We call this mismatch the verification gap. We evaluate one test-first pipeline on selected HumanEval and LiveCodeBench problems and on RGRBench, a collection of 89 Python packages. With one inexpensive model writing tests, generating code, and reviewing both, the withheld tests reject 18.5% of claimed successes on HumanEval, 28.9% on LiveCodeBench, and 41.5% on RGRBench. We then test stronger models as test authors and reviewers while keeping the code generator fixed. Across three providers on the single-function benchmarks, some settings produce fewer success claims that fail the withheld tests, but the evidence for a reduction remains inconclusive. In our main LiveCodeBench experiment, a stronger reviewer also helps the pipeline produce more solutions that pass the withheld tests. On RGRBench, the same reviewer stops many tasks before code is written and produces fewer passing solutions. Failures remain among claimed successes in every configuration studied. The practical lesson is to measure both what an agent produces and how reliably it judges its own work. That judgment also depends on a clear task: some benchmark failures reflect requirements that omit details demanded by the tests. We release the pipeline, recorded runs, and analysis tools to support independent checks of coding agents' success claims.

Zenodo (CERN European Organization for Nuclear Research)
Microsoft (United States) (US), Integris Health (US)
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Who Grades the Grader? The Verification Gap in Self-Testing Code Agents — Brian Sam-Bodden, Eric Kasper, et al. · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS