Who Grades the Grader? The Verification Gap in Self-Testing Code Agents
When a coding agent writes both code and tests, the two can agree while sharing the same mistake. Passing those tests may then conceal a misunderstood requirement. We measure how often code reported as successful fails tests that were not provided to the agent. We call this mismatch the verification gap. We evaluate one test-first pipeline on selected HumanEval and LiveCodeBench problems and on RGRBench, a collection of 89 Python packages. With one inexpensive model writing tests, generating code, and reviewing both, the withheld tests reject 18.5% of claimed successes on HumanEval, 28.9% on LiveCodeBench, and 41.5% on RGRBench. We then test stronger models as test authors and reviewers while keeping the code generator fixed. Across three providers on the single-function benchmarks, some settings produce fewer success claims that fail the withheld tests, but the evidence for a reduction remains inconclusive. In our main LiveCodeBench experiment, a stronger reviewer also helps the pipeline produce more solutions that pass the withheld tests. On RGRBench, the same reviewer stops many tasks before code is written and produces fewer passing solutions. Failures remain among claimed successes in every configuration studied. The practical lesson is to measure both what an agent produces and how reliably it judges its own work. That judgment also depends on a clear task: some benchmark failures reflect requirements that omit details demanded by the tests. We release the pipeline, recorded runs, and analysis tools to support independent checks of coding agents' success claims.
Authors
- Brian Sam-Bodden
- Eric Kasper
- Josh Meyers
- Jyoti Vasudev
Institutions
- Microsoft (United States) (US)
- Integris Health (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23226464
- Primary Topic
- Software Engineering Research
- Type
- preprint