ClerkBench: Measuring the Consistency Horizon of Language Models on Long Deterministic Workloads
Frontier language models can be highly accurate on a task yet fail to perform it reliably enough to run unattended. On long chains of simple, fully specified clerical computation, such as billing, payroll and ledger replay, the binding constraint is not capability but consistency, and ClerkBench measures both as workload lengths. The capability horizon is the length at which a single attempt's probability of producing a byte-exact artifact falls to 50%; the headline consistency horizon is the length at which succeeding on all six of six attempts falls to 50%. Models are tested without tools, so the measurement is the model's own execution reliability. Each instance is generated from a seed and graded against a simulator; the scored set is frozen and released in full. Across fourteen configurations from four frontier laboratories, horizons run from below the smallest rung (50 events) to beyond the largest (400). GPT-6 Astra and Claude Opus 5.5 outrun the ladder on at least one family, so their scores are lower bounds (≥400 and ≥317). Among located horizons GPT-5.5 leads at 147 events, ahead of DeepSeek V4.1 Flash at 138 for about a tenth of the price per attempt, but the leading scores' bootstrap intervals overlap. Matched controls rule out a pure copy or serialization bottleneck: the horizon measures computation, though the score is operational rather than mechanistic. Repeating each task twenty times separates two kinds of failure: random slips, whose consistency cost follows from single-run accuracy, and model–instance traps, where a model fails one instance 16 times in 20 while near-identical siblings pass 17 and 19 of 20. Traps are invisible to single-run scores and did not transfer between the models tested, so consistency must be measured, not inferred. Code and data: github.com/adamallcock/clerkbench, archived as doi:10.5281/zenodo.23048121.
Authors
- Adam Allcock (ORCID: https://orcid.org/0009-0004-0209-8758)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23048151
- Primary Topic
- Natural Language Processing Techniques
- Type
- preprint