Which AI model is truly the best: how reliable are public benchmarks of frontier AI models?

Public leaderboards order AI models by a point score and the public debate about them turns on leads of one or two points. We ask how many of those leads the benchmarks can detect. Using only the item-level results that the benchmarks themselves publish, we computed a psychometric report card for 61 public leaderboards drawn from HELM, the Desai et al. release, SWE-bench Verified, tau2-bench and the archived Open LLM Leaderboard, applying one instrument throughout: exact paired tests on shared items with false-discovery control, item-sampling error, and a tier count down each ranking. 55 of the 58 leaderboards that discriminate at all cannot fully separate their own top five models, 16 separate none of those ten pairs, and in 40 the leader cannot be told apart from at least one other model. On the benchmarks that decide current headlines, Gemini 3 Pro and GPT-5 are statistically tied on GPQA; on SWE-bench Verified, with the agent held fixed, Claude Opus 4.5 and 4.6 are indistinguishable; among the newest models on tau2-bench no named head-to-head is a real difference. Precision follows the square-root law in the number of items (slope -0.49), and a cohort 84 times larger does not improve resolution. We also find that a change of harness build moves one model's tau2-bench score by 14.9 points, more than the gaps it is used to rank, and that published leaderboard data carry defects that can change a model's standing. We recommend that leaderboards report ties and error bars, fix and disclose their harness, and size their benchmarks to the differences they are used to adjudicate.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23007588
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Which AI model is truly the best: how reliable are public benchmarks of frontier AI models?

Miloš Kankaraš
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
preprint

Which AI model is truly the best: how reliable are public benchmarks of frontier AI models?

Miloš Kankaraš
preprint en

Abstract

Public leaderboards order AI models by a point score and the public debate about them turns on leads of one or two points. We ask how many of those leads the benchmarks can detect. Using only the item-level results that the benchmarks themselves publish, we computed a psychometric report card for 61 public leaderboards drawn from HELM, the Desai et al. release, SWE-bench Verified, tau2-bench and the archived Open LLM Leaderboard, applying one instrument throughout: exact paired tests on shared items with false-discovery control, item-sampling error, and a tier count down each ranking. 55 of the 58 leaderboards that discriminate at all cannot fully separate their own top five models, 16 separate none of those ten pairs, and in 40 the leader cannot be told apart from at least one other model. On the benchmarks that decide current headlines, Gemini 3 Pro and GPT-5 are statistically tied on GPQA; on SWE-bench Verified, with the agent held fixed, Claude Opus 4.5 and 4.6 are indistinguishable; among the newest models on tau2-bench no named head-to-head is a real difference. Precision follows the square-root law in the number of items (slope -0.49), and a cohort 84 times larger does not improve resolution. We also find that a change of harness build moves one model's tau2-bench score by 14.9 points, more than the gaps it is used to rank, and that published leaderboard data carry defects that can change a model's standing. We recommend that leaderboards report ties and error bars, fix and disclose their harness, and size their benchmarks to the differences they are used to adjudicate.

Zenodo (CERN European Organization for Nuclear Research)
Reduced inequalities
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.