How Strong Is Your Validator? Measuring Acceptance Gates for LLM-Synthesized Programs

Architectures that let an LLM synthesize disposable programs and accept their outputs through an immutable validator are only as strong as the validator. A previous study of such an architecture (IDAE) showed that format-level validators let well-formed but wrong outputs through, and that one faulty program then corrupts every record of its format. This paper asks how to build a stronger validator and how to measure its strength before deployment. We assemble a corpus of 1,068 unique programs actually synthesized by Claude Opus 4.5 and Claude Haiku 4.5 for three data-extraction domains with a per-record ground-truth oracle; 70 of them are defective, with 23 distinct faulty behaviors (misread number conventions, time-zone and epoch errors, guessed values for missing fields, fabricated fields under an unsatisfiable specification). We compare format validators, source-anchored validators, property validators that relate the output to the raw input, metamorphic validators, N-version agreement, and 80 validators written by the LLMs themselves. Property validators written from the specification alone, without access to the corpus, caught 21 of the 23 faulty behaviors (format validators: 6), rejecting 0.4% of correct outputs. N-version agreement caught 78–100% but rejected 8–27% of correct outputs. LLM-written validators traded silent corruption for false rejection: only 1 of 10 Opus 4.5 and 0 of 10 Haiku 4.5 validators kept false rejection at or below 1% in the financial domain. Letting the same models revise their validators for up to five rounds against automatic self-checks on the specification's examples did not close the gap: in the financial and IoT domains, 4 of 40 validators were usable before and 4 of 40 after, because validators tuned to one example per format still rejected correct records the examples did not cover. A gate that requires each new program to reproduce a few verified input–output pairs of its format never rejected a correct program and caught exactly the faults the property validators missed; combined, they caught 96–100% of the faulty behaviors. A cheap mutation score correlated with strength against real faults (Spearman ρ = 0.51, 95% CI [0.36, 0.66], 100 validators), useful as a screen but not as a substitute. The two faults that escaped the property validators were a plausible re-conversion and an ambiguity in the specification itself, which no validator derived from it can resolve, but a verified example can. English version of the article originally written in Portuguese (https://doi.org/10.5281/zenodo.23074364). Follow-up to IDAE (https://doi.org/10.5281/zenodo.23062292). Code, data, frozen validators and logs of every LLM call: https://github.com/carloseduardodb/idae-validators. Version 2: adds RQ6 (task impact, experiment X7): 14 consumption tasks frozen before the analysis; 20 of the 23 faulty behaviors change the result of some task, and the format validator catches precisely the three harmless ones.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-01
DOI
https://doi.org/10.5281/zenodo.23074532
Primary Topic
Software Engineering Research
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

How Strong Is Your Validator? Measuring Acceptance Gates for LLM-Synthesized Programs

Carlos Eduardo Dias Batista
Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
preprint

How Strong Is Your Validator? Measuring Acceptance Gates for LLM-Synthesized Programs

Carlos Eduardo Dias Batista
preprint en

Abstract

Architectures that let an LLM synthesize disposable programs and accept their outputs through an immutable validator are only as strong as the validator. A previous study of such an architecture (IDAE) showed that format-level validators let well-formed but wrong outputs through, and that one faulty program then corrupts every record of its format. This paper asks how to build a stronger validator and how to measure its strength before deployment. We assemble a corpus of 1,068 unique programs actually synthesized by Claude Opus 4.5 and Claude Haiku 4.5 for three data-extraction domains with a per-record ground-truth oracle; 70 of them are defective, with 23 distinct faulty behaviors (misread number conventions, time-zone and epoch errors, guessed values for missing fields, fabricated fields under an unsatisfiable specification). We compare format validators, source-anchored validators, property validators that relate the output to the raw input, metamorphic validators, N-version agreement, and 80 validators written by the LLMs themselves. Property validators written from the specification alone, without access to the corpus, caught 21 of the 23 faulty behaviors (format validators: 6), rejecting 0.4% of correct outputs. N-version agreement caught 78–100% but rejected 8–27% of correct outputs. LLM-written validators traded silent corruption for false rejection: only 1 of 10 Opus 4.5 and 0 of 10 Haiku 4.5 validators kept false rejection at or below 1% in the financial domain. Letting the same models revise their validators for up to five rounds against automatic self-checks on the specification's examples did not close the gap: in the financial and IoT domains, 4 of 40 validators were usable before and 4 of 40 after, because validators tuned to one example per format still rejected correct records the examples did not cover. A gate that requires each new program to reproduce a few verified input–output pairs of its format never rejected a correct program and caught exactly the faults the property validators missed; combined, they caught 96–100% of the faulty behaviors. A cheap mutation score correlated with strength against real faults (Spearman ρ = 0.51, 95% CI [0.36, 0.66], 100 validators), useful as a screen but not as a substitute. The two faults that escaped the property validators were a plausible re-conversion and an ambiguity in the specification itself, which no validator derived from it can resolve, but a verified example can. English version of the article originally written in Portuguese (https://doi.org/10.5281/zenodo.23074364). Follow-up to IDAE (https://doi.org/10.5281/zenodo.23062292). Code, data, frozen validators and logs of every LLM call: https://github.com/carloseduardodb/idae-validators. Version 2: adds RQ6 (task impact, experiment X7): 14 consumption tasks frozen before the analysis; 20 of the 23 faulty behaviors change the result of some task, and the format validator catches precisely the three harmless ones.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.