Verify-and-Repair Under Matched Call Budgets: Resampling, Grammar Constraints, and Memory Wiring
Wrappers that add a verify-and-repair loop around a language model are widely reported to improve instruction following. We ask two narrower questions. Does the loop show an advantage over verifier-selected resampling when both are allowed the same number of generations? Does wiring the loop's failures into a durable memory show an advantage over the same components left unwired? Each question was tested twice on machine-checkable format constraints, inside a set of six studies whose plans were frozen by hash before any call; two of the six were also registered publicly on OSF before their runs. Under the decision rule frozen in each plan, none of the six primary comparisons separated from zero. For repair against resampling the point estimates were +6.2 and +1.5 points. For the wired stack against the unwired one they were -2.8 and -6.1 points. Grammar-constrained decoding met the format on every accepted, completed generation under the tested grammars, with three of 39 grammars rejected by the server's parser and counted as failures under the frozen rule. In the second memory study, three lesson policies raised first-attempt pass rates, while the share of first-attempt failures recovered by retries was lower than in the unwired arm. Each arm's retry set is its own set of failures, so that pattern does not by itself show that the lessons impaired the retries. The same tasks and checkers registered large specified differences elsewhere, from 11 to 97 points, which shows the instrument was responsive without establishing power for every primary comparison. We report the intervals, the candidate explanations the records support, the registration practice we would change, and two defects found in the IFEval reference checkers. Preregistrations: 10.17605/OSF.IO/QC3DB (E8) and 10.17605/OSF.IO/T6JKZ (E9). Frozen plans, raw model calls and analysis scripts: the reviewer archive at 10.17605/OSF.IO/EDUAK and github.com/oceanusgascon26/blue-fairy-reviewer-archive. Evaluation kit: 10.5281/zenodo.22683481. 12 pages, 2 figures, 4 tables.
Authors
- Chris Gascon
Institutions
- Ocean Networks Canada Society (CA)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-17
- DOI
- https://doi.org/10.5281/zenodo.22817238
- Primary Topic
- Natural Language Processing Techniques
- Type
- preprint