Verify-and-Repair Under Matched Call Budgets: Resampling, Grammar Constraints, and Memory Wiring

Wrappers that add a verify-and-repair loop around a language model are widely reported to improve instruction following. We ask two narrower questions. Does the loop show an advantage over verifier-selected resampling when both are allowed the same number of generations? Does wiring the loop's failures into a durable memory show an advantage over the same components left unwired? Each question was tested twice on machine-checkable format constraints, inside a set of six studies whose plans were frozen by hash before any call; two of the six were also registered publicly on OSF before their runs. Under the decision rule frozen in each plan, none of the six primary comparisons separated from zero. For repair against resampling the point estimates were +6.2 and +1.5 points. For the wired stack against the unwired one they were -2.8 and -6.1 points. Grammar-constrained decoding met the format on every accepted, completed generation under the tested grammars, with three of 39 grammars rejected by the server's parser and counted as failures under the frozen rule. In the second memory study, three lesson policies raised first-attempt pass rates, while the share of first-attempt failures recovered by retries was lower than in the unwired arm. Each arm's retry set is its own set of failures, so that pattern does not by itself show that the lessons impaired the retries. The same tasks and checkers registered large specified differences elsewhere, from 11 to 97 points, which shows the instrument was responsive without establishing power for every primary comparison. We report the intervals, the candidate explanations the records support, the registration practice we would change, and two defects found in the IFEval reference checkers. Preregistrations: 10.17605/OSF.IO/QC3DB (E8) and 10.17605/OSF.IO/T6JKZ (E9). Frozen plans, raw model calls and analysis scripts: the reviewer archive at 10.17605/OSF.IO/EDUAK and github.com/oceanusgascon26/blue-fairy-reviewer-archive. Evaluation kit: 10.5281/zenodo.22683481. 12 pages, 2 figures, 4 tables.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-17
DOI
https://doi.org/10.5281/zenodo.22817239
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Verify-and-Repair Under Matched Call Budgets: Resampling, Grammar Constraints, and Memory Wiring

Chris Gascon
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

Verify-and-Repair Under Matched Call Budgets: Resampling, Grammar Constraints, and Memory Wiring

Chris Gascon
preprint en

Abstract

Wrappers that add a verify-and-repair loop around a language model are widely reported to improve instruction following. We ask two narrower questions. Does the loop show an advantage over verifier-selected resampling when both are allowed the same number of generations? Does wiring the loop's failures into a durable memory show an advantage over the same components left unwired? Each question was tested twice on machine-checkable format constraints, inside a set of six studies whose plans were frozen by hash before any call; two of the six were also registered publicly on OSF before their runs. Under the decision rule frozen in each plan, none of the six primary comparisons separated from zero. For repair against resampling the point estimates were +6.2 and +1.5 points. For the wired stack against the unwired one they were -2.8 and -6.1 points. Grammar-constrained decoding met the format on every accepted, completed generation under the tested grammars, with three of 39 grammars rejected by the server's parser and counted as failures under the frozen rule. In the second memory study, three lesson policies raised first-attempt pass rates, while the share of first-attempt failures recovered by retries was lower than in the unwired arm. Each arm's retry set is its own set of failures, so that pattern does not by itself show that the lessons impaired the retries. The same tasks and checkers registered large specified differences elsewhere, from 11 to 97 points, which shows the instrument was responsive without establishing power for every primary comparison. We report the intervals, the candidate explanations the records support, the registration practice we would change, and two defects found in the IFEval reference checkers. Preregistrations: 10.17605/OSF.IO/QC3DB (E8) and 10.17605/OSF.IO/T6JKZ (E9). Frozen plans, raw model calls and analysis scripts: the reviewer archive at 10.17605/OSF.IO/EDUAK and github.com/oceanusgascon26/blue-fairy-reviewer-archive. Evaluation kit: 10.5281/zenodo.22683481. 12 pages, 2 figures, 4 tables.

Zenodo (CERN European Organization for Nuclear Research)
Ocean Networks Canada Society (CA)
Quality Education
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.