How Much of a Leaderboard Ranking Survives Its Own Sampling Error? A Paired-Resolution Audit of Five AI Benchmarks
A benchmark leaderboard is read as a stopwatch: the rank beside a model is treated as a fact. Yet every entrant is scored on the same items and publishes which it solved, so a leaderboard is a paired experiment, and its ordering can be tested rather than trusted. We test each adjacent-rank pair for statistical separability - exact McNemar's test on discordant items for pass/fail benchmarks, a paired item-bootstrap for scored ones - across five public boards. On SWE-bench Verified, 129 of 133 adjacent ranks are not separable at the 5% level; on MTEB, 176 of 180; on HELM Lite, all 89 of 89. The benchmarks are not broken: 88% of all pairs separate cleanly. It is the neighbours they cannot order, and neighbours are what the top of a table is made of. We derive a resolution bound, delta >= 2.80*sqrt(d/n) in the discordance rate d, that predicts which pairs are decidable from a benchmark's size, validate it on 12,739 pairs with zero false positives, and use LMArena - which already ships confidence intervals and ties - as a positive control for what correct reporting looks like. Every figure is recomputed from public data and passes a self-test suite; the recomputed rates match each official board exactly. We argue that a board should report the resolution of its ranking alongside the ranking, at no additional data cost.
Authors
- Ilpo Väätäinen
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-06
- DOI
- https://doi.org/10.5281/zenodo.22541066
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint