When Reviewers Agree on the Wrong Reading: A Pilot Study of Validator Independence and Specification Ambiguity in the Assurance of Agent-Written Code
Organizations that let AI agents change production code are increasingly relying on a chain of evidence before accepting a change: the agent's own tests, a second agent's review, a different model's judgment. We ask how often the organization makes the correct accept, reject or hold decision given that chain, and hypothesize that when the requirement is ambiguous, the ambiguity acts as a common cause that makes independent reviewers fail together. We report a small, fully disclosed pilot: five insurance-domain tasks with a planted ambiguity, implemented by gpt-5.1, with wrong and correct implementations judged at four independence levels by gpt-5.1, Claude Haiku 5.5 and Claude Sonnet 5.5. Within the same task, no validator configuration reliably accepted correct code more often than wrong code. The pilot has five tasks, one implementer family and single-run cells, so it supports a hypothesis and a measurement method, not a general claim. Code, tasks and raw verdicts are released with a DOI.
Authors
- Moulinath Chakrabarty (ORCID: https://orcid.org/0009-0002-3084-2930)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23234729
- Primary Topic
- Software Engineering Research
- Type
- preprint