ClerkBench: Measuring the Consistency Horizon of Language Models on Long Deterministic Workloads

Frontier language models can be highly accurate on a task yet fail to perform it reliably enough to run unattended. On long chains of simple, fully specified clerical computation, such as billing, payroll and ledger replay, the binding constraint is not capability but consistency, and ClerkBench measures both as workload lengths. The capability horizon is the length at which a single attempt's probability of producing a byte-exact artifact falls to 50%; the headline consistency horizon is the length at which succeeding on all six of six attempts falls to 50%. Models are tested without tools, so the measurement is the model's own execution reliability. Each instance is generated from a seed and graded against a simulator; the scored set is frozen and released in full. Across fourteen configurations from four frontier laboratories, horizons run from below the smallest rung (50 events) to beyond the largest (400). GPT-6 Astra and Claude Opus 5.5 outrun the ladder on at least one family, so their scores are lower bounds (≥400 and ≥317). Among located horizons GPT-5.5 leads at 147 events, ahead of DeepSeek V4.1 Flash at 138 for about a tenth of the price per attempt, but the leading scores' bootstrap intervals overlap. Matched controls rule out a pure copy or serialization bottleneck: the horizon measures computation, though the score is operational rather than mechanistic. Repeating each task twenty times separates two kinds of failure: random slips, whose consistency cost follows from single-run accuracy, and model–instance traps, where a model fails one instance 16 times in 20 while near-identical siblings pass 17 and 19 of 20. Traps are invisible to single-run scores and did not transfer between the models tested, so consistency must be measured, not inferred. Code and data: github.com/adamallcock/clerkbench, archived as doi:10.5281/zenodo.23048121.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23048150
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

ClerkBench: Measuring the Consistency Horizon of Language Models on Long Deterministic Workloads

Adam Allcock
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

ClerkBench: Measuring the Consistency Horizon of Language Models on Long Deterministic Workloads

Adam Allcock
preprint en

Abstract

Frontier language models can be highly accurate on a task yet fail to perform it reliably enough to run unattended. On long chains of simple, fully specified clerical computation, such as billing, payroll and ledger replay, the binding constraint is not capability but consistency, and ClerkBench measures both as workload lengths. The capability horizon is the length at which a single attempt's probability of producing a byte-exact artifact falls to 50%; the headline consistency horizon is the length at which succeeding on all six of six attempts falls to 50%. Models are tested without tools, so the measurement is the model's own execution reliability. Each instance is generated from a seed and graded against a simulator; the scored set is frozen and released in full. Across fourteen configurations from four frontier laboratories, horizons run from below the smallest rung (50 events) to beyond the largest (400). GPT-6 Astra and Claude Opus 5.5 outrun the ladder on at least one family, so their scores are lower bounds (≥400 and ≥317). Among located horizons GPT-5.5 leads at 147 events, ahead of DeepSeek V4.1 Flash at 138 for about a tenth of the price per attempt, but the leading scores' bootstrap intervals overlap. Matched controls rule out a pure copy or serialization bottleneck: the horizon measures computation, though the score is operational rather than mechanistic. Repeating each task twenty times separates two kinds of failure: random slips, whose consistency cost follows from single-run accuracy, and model–instance traps, where a model fails one instance 16 times in 20 while near-identical siblings pass 17 and 19 of 20. Traps are invisible to single-run scores and did not transfer between the models tested, so consistency must be measured, not inferred. Code and data: github.com/adamallcock/clerkbench, archived as doi:10.5281/zenodo.23048121.

Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.