How Much of a Leaderboard Ranking Survives Its Own Sampling Error? A Paired-Resolution Audit of Five AI Benchmarks

{"A":[0],"benchmark":[1,552],"leaderboard":[2,34],"is":[3,13,21,35,116,131,190,387,436,582,615,692],"read":[4],"as":[5,15,178,349,647,649,679],"a":[6,11,16,33,36,66,129,136,154,179,197,220,225,231,292,330,439,460,518,536,575,594,682,712],"stopwatch:":[7],"the":[8,24,89,117,126,143,200,215,235,241,262,268,272,350,354,363,390,423,430,466,487,490,494,503,513,529,540,545,549,572,583,585,611,627,634,652,677,699,719,723],"rank":[9,360],"beside":[10],"model":[12,303],"treated":[14],"fact.":[17],"Yet":[18],"every":[19,256],"entrant":[20],"scored":[22,70],"on":[23,60,92,97,159,214,512,630,636,640,715],"same":[25,431,707],"items":[26,62,479],"and":[27,39,122,166,175,195,265,275,309,321,327,342,357,370,378,404,417,474,544,559,578,604,618,639,651,698,718,738],"publishes":[28],"which":[29,149,170,396,581,655,732],"it":[30,158,435,463,691,729],"solved,":[31],"so":[32,643],"paired":[37,67],"experiment,":[38],"its":[40,208,238,374,398,405],"ordering":[41],"can":[42,298],"be":[43],"tested":[44],"rather":[45,333,555],"than":[46,334,556],"trusted.":[47],"We":[48,134,228],"test":[49,59],"each":[50,204,498],"adjacent-rank":[51],"pair":[52],"for":[53,63,69,182,353,358,400],"statistical":[54],"separability":[55],"-":[56,72,169,177],"exact":[57],"McNemar's":[58],"discordant":[61],"pass/fail":[64],"benchmarks,":[65],"item-bootstrap":[68],"ones":[71],"across":[73],"five":[74],"public":[75,193],"boards.":[76],"On":[77],"SWE-bench":[78],"Verified,":[79],"129":[80],"of":[81,95,102,110,128,237,295,392,434,462,509,633,684,735,740],"133":[82],"adjacent":[83],"ranks":[84],"are":[85,106,124,151,282,346,427],"not":[86,107,311,739],"separable":[87],"at":[88,243,410],"5%":[90],"level;":[91],"MTEB,":[93],"176":[94],"180;":[96],"HELM":[98],"Lite,":[99],"all":[100,111,276,401,703],"89":[101],"89.":[103],"The":[104,288,562,671,706],"benchmarks":[105],"broken:":[108],"88%":[109],"pairs":[112,150,161],"separate":[113],"cleanly.":[114],"It":[115],"neighbours":[118,123],"they":[119],"cannot":[120],"order,":[121],"what":[125,183],"top":[127,302],"table":[130,541],"made":[132],"of.":[133],"derive":[135],"resolution":[137,236],"bound,":[138],"delta":[139],">=":[140],"2.80*sqrt(d/n)":[141],"in":[142,224,429,438,443,471,476,535,568,610,711],"discordance":[144],"rate":[145],"d,":[146],"that":[147,218,230,290,465,726],"predicts":[148],"decidable":[152],"from":[153,192,261,539,551,557],"benchmark's":[155],"size,":[156],"validate":[157],"12,739":[160],"with":[162,211],"zero":[163],"false":[164],"positives,":[165],"use":[167],"LMArena":[168],"already":[171],"ships":[172],"confidence":[173,361],"intervals":[174],"ties":[176],"positive":[180],"control":[181],"correct":[184],"reporting":[185],"looks":[186],"like.":[187],"Every":[188],"figure":[189,278],"recomputed":[191,201,495],"data":[194,246,264,297,323],"passes":[196],"self-test":[198],"suite;":[199],"rates":[202,496],"match":[203,497],"official":[205,499],"board":[206,232,500],"to":[207,305,312,365,414,517,588,722],"published":[209,384],"precision,":[210],"one":[212,388,472,507,736],"mismatch":[213,508],"Test":[216,514,631],"split":[217],"reproduces":[219,397],"known":[221],"duplicate-id":[222],"defect":[223,584],"submission":[226,519],"file.":[227],"argue":[229],"should":[233],"report":[234],"ranking":[239,385,399],"alongside":[240],"ranking,":[242],"no":[244,579,607,657],"additional":[245],"cost.":[247],"Version":[248,446,451,664],"2":[249,413,452,570,592],"(18":[250,448],"September":[251,449,667],"2026).":[252,450,668],"No":[253],"measurement":[254],"changed:":[255],"quantitative":[257],"result":[258,289],"was":[259,645,709,730,733],"rerun":[260],"deposited":[263,621],"reproduced,":[266],"including":[267],"three":[269],"parity":[270],"gates,":[271],"preregistration":[273],"scoring":[274],"four":[277,704],"coordinates.":[279],"What":[280],"changed":[281],"sentences":[283],"about":[284,422,456],"other":[285,457],"people's":[286,458],"work.":[287],"dropping":[291],"small":[293],"share":[294],"preference":[296],"change":[299],"Chatbot":[300],"Arena's":[301],"belongs":[304],"Huang,":[306],"Shen,":[307],"Wei":[308],"Broderick,":[310],"Singh":[313],"et":[314,339,344,564,673],"al.,":[315],"whose":[316,600],"paper":[317,488,714],"documents":[318],"private":[319],"testing":[320],"unequal":[322],"access":[324],"instead;":[325],"Zhang":[326],"Hardt":[328],"measure":[329],"diversity-stability":[331],"trade-off":[332],"proving":[335],"an":[336],"impossibility;":[337],"Card":[338],"al.":[340,345,565,674],"(2020)":[341],"Mogstad":[343,563],"now":[347,616,660,701],"cited":[348],"prior":[351],"art":[352],"power":[355],"requirement":[356],"simultaneous":[359],"sets;":[362],"citation":[364],"Lunardi,":[366],"Della":[367],"Mea,":[368],"Mizzaro":[369],"Roitero":[371],"named":[372],"only":[373],"last":[375],"two":[376],"authors":[377],"had":[379,468,606],"strengthened":[380],"their":[381],"hedges;":[382],"LMArena's":[383],"rule":[386],"plus":[389],"number":[391],"models":[393,408],"significantly":[394],"above,":[395],"242":[402],"models,":[403],"ten":[406],"rank-2":[407],"sit":[409],"rating":[411],"positions":[412],"9,":[415],"11":[416],"20.":[418],"Four":[419],"checkable":[420],"claims":[421],"author's":[424],"own":[425],"artefacts":[426],"corrected":[428,453,710],"pass.":[432],"All":[433],"listed":[437,534],"new":[440],"section,":[441],"\\"Corrections":[442],"this":[444,741],"version\\".":[445],"3":[447],"twelve":[454],"statements":[455],"work;":[459],"re-audit":[461],"found":[464],"correction":[467,586],"been":[469],"incomplete":[470],"direction":[473],"wrong":[475,573,646],"another.":[477],"Three":[478],"reached":[480],"version":[481,569],"2's":[482],"deposit":[483,737],"description":[484],"but":[485],"never":[486],"itself:":[489],"abstract":[491,546],"still":[492,547],"claimed":[493],"exactly,":[501],"where":[502],"results":[504,523],"file":[505,520],"records":[506],"1.22":[510],"pp":[511],"split,":[515,654],"traceable":[516],"listing":[521],"241":[522],"over":[524,623],"213":[525],"distinct":[526],"instance":[527],"ids;":[528],"per-split":[530],"agreement":[531],"figures":[532],"were":[533],"different":[537],"order":[538],"above":[542],"them;":[543],"derived":[548],"bound":[550],"size":[553,558],"alone":[554],"near-pair":[560],"discordance.":[561],"entry":[566,675,700],"added":[567],"carried":[571,593],"year,":[574],"truncated":[576],"title":[577,678],"venue,":[580],"existed":[587],"remove.":[589],"And":[590],"Table":[591],"column":[595],"headed":[596],"\\"#1":[597],"survives":[598],"reshuffle\\"":[599],"values,":[601],"\\"every":[602],"time\\"":[603],"\\"never\\",":[605],"computation":[608],"anywhere":[609],"repository.":[612],"That":[613],"check":[614],"defined":[617],"run":[619],"(top1_stability.py,":[620],"here):":[622],"2,000":[624],"item":[625],"bootstraps":[626],"leader":[628],"holds":[629],"100%":[632],"time,":[635],"Verified":[637],"38.3%":[638],"Lite":[641],"53.6%,":[642],"\\"never\\"":[644],"well":[648],"unevidenced,":[650],"Multimodal":[653],"has":[656],"per-instance":[658],"matrix,":[659],"reads":[661],"\\"not":[662],"computed\\".":[663],"4":[665],"(19":[666],"One":[669],"reference.":[670],"Huang":[672],"gave":[676],"\\"Dropping":[680],"Just":[681],"Handful":[683],"Preferences":[685],"Can":[686],"Change":[687],"Top":[688],"LLM":[689],"Rankings\\";":[690],"\\"Top":[693],"Large":[694],"Language":[695],"Model":[696],"Rankings\\",":[697],"names":[702],"authors.":[705],"truncation":[708],"companion":[713],"18":[716],"September,":[717],"letter":[720],"sent":[721],"first":[724],"author":[725],"day":[727],"said":[728],"fixed,":[731],"true":[734],"one.":[742],"Nothing":[743],"else":[744],"changed.":[745]}

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-19
DOI
https://doi.org/10.5281/zenodo.22839458
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

How Much of a Leaderboard Ranking Survives Its Own Sampling Error? A Paired-Resolution Audit of Five AI Benchmarks

Ilpo Väätäinen
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

How Much of a Leaderboard Ranking Survives Its Own Sampling Error? A Paired-Resolution Audit of Five AI Benchmarks

Ilpo Väätäinen
preprint en

Abstract

A benchmark leaderboard is read as a stopwatch: the rank beside a model is treated as a fact. Yet every entrant is scored on the same items and publishes which it solved, so a leaderboard is a paired experiment, and its ordering can be tested rather than trusted. We test each adjacent-rank pair for statistical separability - exact McNemar's test on discordant items for pass/fail benchmarks, a paired item-bootstrap for scored ones - across five public boards. On SWE-bench Verified, 129 of 133 adjacent ranks are not separable at the 5% level; on MTEB, 176 of 180; on HELM Lite, all 89 of 89. The benchmarks are not broken: 88% of all pairs separate cleanly. It is the neighbours they cannot order, and neighbours are what the top of a table is made of. We derive a resolution bound, delta >= 2.80*sqrt(d/n) in the discordance rate d, that predicts which pairs are decidable from a benchmark's size, validate it on 12,739 pairs with zero false positives, and use LMArena - which already ships confidence intervals and ties - as a positive control for what correct reporting looks like. Every figure is recomputed from public data and passes a self-test suite; the recomputed rates match each official board to its published precision, with one mismatch on the Test split that reproduces a known duplicate-id defect in a submission file. We argue that a board should report the resolution of its ranking alongside the ranking, at no additional data cost. Version 2 (18 September 2026). No measurement changed: every quantitative result was rerun from the deposited data and reproduced, including the three parity gates, the preregistration scoring and all four figure coordinates. What changed are sentences about other people's work. The result that dropping a small share of preference data can change Chatbot Arena's top model belongs to Huang, Shen, Wei and Broderick, not to Singh et al., whose paper documents private testing and unequal data access instead; Zhang and Hardt measure a diversity-stability trade-off rather than proving an impossibility; Card et al. (2020) and Mogstad et al. are now cited as the prior art for the power requirement and for simultaneous rank confidence sets; the citation to Lunardi, Della Mea, Mizzaro and Roitero named only its last two authors and had strengthened their hedges; LMArena's published ranking rule is one plus the number of models significantly above, which reproduces its ranking for all 242 models, and its ten rank-2 models sit at rating positions 2 to 9, 11 and 20. Four checkable claims about the author's own artefacts are corrected in the same pass. All of it is listed in a new section, "Corrections in this version". Version 3 (18 September 2026). Version 2 corrected twelve statements about other people's work; a re-audit of it found that the correction had been incomplete in one direction and wrong in another. Three items reached version 2's deposit description but never the paper itself: the abstract still claimed the recomputed rates match each official board exactly, where the results file records one mismatch of 1.22 pp on the Test split, traceable to a submission file listing 241 results over 213 distinct instance ids; the per-split agreement figures were listed in a different order from the table above them; and the abstract still derived the bound from benchmark size alone rather than from size and near-pair discordance. The Mogstad et al. entry added in version 2 carried the wrong year, a truncated title and no venue, which is the defect the correction existed to remove. And Table 2 carried a column headed "#1 survives reshuffle" whose values, "every time" and "never", had no computation anywhere in the repository. That check is now defined and run (top1_stability.py, deposited here): over 2,000 item bootstraps the leader holds on Test 100% of the time, on Verified 38.3% and on Lite 53.6%, so "never" was wrong as well as unevidenced, and the Multimodal split, which has no per-instance matrix, now reads "not computed". Version 4 (19 September 2026). One reference. The Huang et al. entry gave the title as "Dropping Just a Handful of Preferences Can Change Top LLM Rankings"; it is "Top Large Language Model Rankings", and the entry now names all four authors. The same truncation was corrected in a companion paper on 18 September, and the letter sent to the first author that day said it was fixed, which was true of one deposit and not of this one. Nothing else changed.

Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.