Which AI model is truly the best: how reliable are public benchmarks of frontier AI models?
Public leaderboards order AI models by a point score and the public debate about them turns on leads of one or two points. We ask how many of those leads the benchmarks can detect. Using only the item-level results that the benchmarks themselves publish, we computed a psychometric report card for 61 public leaderboards drawn from HELM, the Desai et al. release, SWE-bench Verified, tau2-bench and the archived Open LLM Leaderboard, applying one instrument throughout: exact paired tests on shared items with false-discovery control, item-sampling error, and a tier count down each ranking. 55 of the 58 leaderboards that discriminate at all cannot fully separate their own top five models, 16 separate none of those ten pairs, and in 40 the leader cannot be told apart from at least one other model. On the benchmarks that decide current headlines, Gemini 3 Pro and GPT-5 are statistically tied on GPQA; on SWE-bench Verified, with the agent held fixed, Claude Opus 4.5 and 4.6 are indistinguishable; among the newest models on tau2-bench no named head-to-head is a real difference. Precision follows the square-root law in the number of items (slope -0.49), and a cohort 84 times larger does not improve resolution. We also find that a change of harness build moves one model's tau2-bench score by 14.9 points, more than the gaps it is used to rank, and that published leaderboard data carry defects that can change a model's standing. We recommend that leaderboards report ties and error bars, fix and disclose their harness, and size their benchmarks to the differences they are used to adjudicate.
Authors
- Miloš Kankaraš (ORCID: https://orcid.org/0000-0002-3190-7751)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23007588
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- preprint