Anytime-Valid Stopping for Self-Consistency on Open Answer Sets
Self-consistency samples reasoning chains and returns the most frequent answer, and adaptive variants stop sampling early to save compute. Deployments act on the stopped answer, yet these rules recheck a fixed-sample criterion after every draw, so the chance of stopping on an answer that is not the population mode has no time-uniform bound: in simulation, Adaptive Consistency (AC) at its published setting exceeds the tolerance implied by its credible level significantly in 27 of 70 i.i.d. cells within its 40-draw cap. Existing anytime-valid mode certificates need the answer set in advance, test a null stronger than modality, or pay a multiplicity charge that grows with the number of answers seen. We give a stopping rule for i.i.d. draws over an unknown, possibly countably infinite answer set that certifies at level δ that an answer it selects from the data is a population mode. Each discovered answer must beat every rival, seen or unseen, against a threshold set by the order in which it was first seen rather than by the size of the answer set. Under a unique mode and runner-up, the expected stopping time of this complete rule attains the optimal leading constant as δ → 0 without knowledge of the answer set. The rule and its four variants never exceed δ in 140 i.i.d. simulation cells; where both certify at least 90%, the guarantee costs a median 3.2–3.4 times AC's draws. On GSM8K with Qwen2.5-1.5B, the rule, set by δ alone, lands within one wrong certificate and four certified questions in 200 of the coverage–error frontier that AC reaches only when tuned in-sample.
Authors
- Ya-Fen Yeh
- Guan-Yuan Chen (ORCID: https://orcid.org/0000-0003-3298-0624)
Institutions
- National Tsing Hua University (TW)
- North Carolina Exploring Cultural Heritage Online (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22959700
- Primary Topic
- Distributed systems and fault tolerance
- Type
- preprint