How Much of a Leaderboard Ranking Survives Its Own Sampling Error? A Paired-Resolution Audit of Five AI Benchmarks

A benchmark leaderboard is read as a stopwatch: the rank beside a model is treated as a fact. Yet every entrant is scored on the same items and publishes which it solved, so a leaderboard is a paired experiment, and its ordering can be tested rather than trusted. We test each adjacent-rank pair for statistical separability - exact McNemar's test on discordant items for pass/fail benchmarks, a paired item-bootstrap for scored ones - across five public boards. On SWE-bench Verified, 129 of 133 adjacent ranks are not separable at the 5% level; on MTEB, 176 of 180; on HELM Lite, all 89 of 89. The benchmarks are not broken: 88% of all pairs separate cleanly. It is the neighbours they cannot order, and neighbours are what the top of a table is made of. We derive a resolution bound, delta >= 2.80*sqrt(d/n) in the discordance rate d, that predicts which pairs are decidable from a benchmark's size, validate it on 12,739 pairs with zero false positives, and use LMArena - which already ships confidence intervals and ties - as a positive control for what correct reporting looks like. Every figure is recomputed from public data and passes a self-test suite; the recomputed rates match each official board exactly. We argue that a board should report the resolution of its ranking alongside the ranking, at no additional data cost.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-06
DOI
https://doi.org/10.5281/zenodo.22541066
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

How Much of a Leaderboard Ranking Survives Its Own Sampling Error? A Paired-Resolution Audit of Five AI Benchmarks

Ilpo Väätäinen
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

How Much of a Leaderboard Ranking Survives Its Own Sampling Error? A Paired-Resolution Audit of Five AI Benchmarks

Ilpo Väätäinen
preprint en

Abstract

A benchmark leaderboard is read as a stopwatch: the rank beside a model is treated as a fact. Yet every entrant is scored on the same items and publishes which it solved, so a leaderboard is a paired experiment, and its ordering can be tested rather than trusted. We test each adjacent-rank pair for statistical separability - exact McNemar's test on discordant items for pass/fail benchmarks, a paired item-bootstrap for scored ones - across five public boards. On SWE-bench Verified, 129 of 133 adjacent ranks are not separable at the 5% level; on MTEB, 176 of 180; on HELM Lite, all 89 of 89. The benchmarks are not broken: 88% of all pairs separate cleanly. It is the neighbours they cannot order, and neighbours are what the top of a table is made of. We derive a resolution bound, delta >= 2.80*sqrt(d/n) in the discordance rate d, that predicts which pairs are decidable from a benchmark's size, validate it on 12,739 pairs with zero false positives, and use LMArena - which already ships confidence intervals and ties - as a positive control for what correct reporting looks like. Every figure is recomputed from public data and passes a self-test suite; the recomputed rates match each official board exactly. We argue that a board should report the resolution of its ranking alongside the ranking, at no additional data cost.

Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.