The Benchmark Number Is Not the Deployment Number: Rank Collapse and Base-Rate Collapse, Measured Across Four Domains

{"A":[0],"leaderboard":[1],"reports":[2],"one":[3,178],"number":[4,10],"per":[5,256],"model.":[6],"Two":[7],"things":[8],"that":[9,179,196],"silently":[11],"discards":[12],"are":[13],"how":[14],"many":[15,266],"independent":[16],"axes":[17,86],"of":[18,70,94,191,241],"skill":[19],"it":[20],"collapses":[21],"into":[22],"a":[23,53,67,91,119,173,194,207,246,276,309,313],"single":[24,79],"scalar,":[25],"and":[26,116,132,153,260,280,312,317],"what":[27],"population":[28,210],"its":[29],"accuracy":[30],"was":[31],"measured":[32],"on":[33,44,61,101,127,134,193,232,272],"versus":[34],"where":[35],"the":[36,63,71,78,140,204,212,236,262,284,289,293,297,318],"model":[37,75],"is":[38,104,197,292],"used.":[39],"This":[40],"paper":[41,302],"measures":[42],"both,":[43],"public":[45],"data,":[46],"with":[47],"every":[48,177],"figure":[49],"independently":[50],"re-verified":[51],"through":[52],"second":[54],"code":[55],"path":[56],"before":[57],"inclusion.":[58],"Rank":[59],"collapse:":[60,184],"ProteinGym,":[62],"standard":[64],"variant-effect-prediction":[65],"benchmark,":[66],"null-gated":[68,157],"decomposition":[69],"217x97":[72],"assay":[73],"x":[74,121],"matrix":[76],"shows":[77,261],"ranking":[80],"hides":[81],"at":[82,181,223,245],"least":[83],"three":[84,161],"specialization":[85],"(singular":[87],"values":[88],"7.87/6.60/5.88":[89],"against":[90],"permutation":[92],"null":[93],"4.46/4.32/4.22).":[95],"The":[96,155,225,301],"sharpest":[97],"axis,":[98],"structure-vs-sequence":[99],"models":[100,125],"protein":[102],"stability,":[103],"powered":[105],"across":[106],"all":[107],"64":[108],"Tsuboyama":[109],"stability":[110],"assays":[111],"(64/64,":[112],"median":[113],"gap":[114],"+0.346)":[115],"traced":[117],"to":[118,160,203,220,229,243],"chemistry":[120],"exposure":[122],"mechanism:":[123],"structure-aware":[124],"win":[126],"surface":[128,135],"aromatic":[129],"substitutions":[130],"(p=8e-8)":[131],"lose":[133],"charge":[136],"substitutions,":[137],"converging":[138],"in":[139,176],"buried":[141],"core,":[142],"after":[143],"two":[144,304],"competing":[145],"hypotheses":[146],"(buried-core":[147],"packing;":[148],"general":[149],"hydrophobicity)":[150],"were":[151],"tested":[152],"killed.":[154],"same":[156,294],"test":[158],"applied":[159,228],"unrelated":[162],"benchmarks":[163],"(MTEB":[164],"embeddings,":[165],"SWE-bench":[166],"code-agent":[167],"evaluation,":[168],"LMArena":[169],"chat":[170],"preference)":[171],"finds":[172],"hidden":[174],"axis":[175],"resolves":[180],"all.":[182],"Base-rate":[183],"ten":[185],"clinical":[186],"variant-pathogenicity":[187],"predictors":[188],"report":[189],"AUCs":[190],"0.81-0.98":[192],"benchmark":[195,239],"65%":[198],"pathogenic;":[199],"reprojected":[200],"by":[201],"Bayes":[202,314],"0.1-1%":[205],"prevalence":[206],"real":[208],"screening":[209],"presents,":[211],"best":[213],"predictor's":[214],"flag":[215],"precision":[216,240],"falls":[217],"from":[218,275],"93%":[219],"12%":[221],"(1.3%":[222],"0.1%).":[224],"identical":[226,285],"arithmetic,":[227],"AI-generated-text":[230],"detection":[231],"GPT-4":[233],"output,":[234],"takes":[235],"strongest":[237],"detector's":[238],"79%":[242],"16%":[244],"5%":[247],"real-world":[248],"machine-text":[249],"rate":[250,299],"--":[251,259,308,316],"about":[252],"five":[253],"false":[254],"accusations":[255],"genuine":[257],"catch":[258],"classifier":[263],"lineage":[264],"behind":[265],"commercial":[267],"\\"AI":[268],"checkers\\"":[269],"scores":[270],"AUC=0.545":[271],"GPT-4,":[273],"indistinguishable":[274],"coin":[277],"flip.":[278],"Genomics":[279],"generated":[281],"text":[282],"share":[283],"failure":[286],"shape":[287],"because":[288],"underlying":[290],"arithmetic":[291],"theorem;":[295],"only":[296],"base":[298],"changes.":[300],"releases":[303],"reusable,":[305],"self-tested":[306],"primitives":[307],"rank-dimensionality":[310],"auditor":[311],"recalibrator":[315],"full":[319],"analysis":[320],"pipeline.":[321]}

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-08-26
DOI
https://doi.org/10.5281/zenodo.22109473
Primary Topic
Genomics and Rare Diseases
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Benchmark Number Is Not the Deployment Number: Rank Collapse and Base-Rate Collapse, Measured Across Four Domains

Ilpo Väätäinen
Zenodo (CERN European Organization for Nuclear Research)
Genomics and Rare Diseases
preprint

The Benchmark Number Is Not the Deployment Number: Rank Collapse and Base-Rate Collapse, Measured Across Four Domains

Ilpo Väätäinen
preprint en

Abstract

A leaderboard reports one number per model. Two things that number silently discards are how many independent axes of skill it collapses into a single scalar, and what population its accuracy was measured on versus where the model is used. This paper measures both, on public data, with every figure independently re-verified through a second code path before inclusion. Rank collapse: on ProteinGym, the standard variant-effect-prediction benchmark, a null-gated decomposition of the 217x97 assay x model matrix shows the single ranking hides at least three specialization axes (singular values 7.87/6.60/5.88 against a permutation null of 4.46/4.32/4.22). The sharpest axis, structure-vs-sequence models on protein stability, is powered across all 64 Tsuboyama stability assays (64/64, median gap +0.346) and traced to a chemistry x exposure mechanism: structure-aware models win on surface aromatic substitutions (p=8e-8) and lose on surface charge substitutions, converging in the buried core, after two competing hypotheses (buried-core packing; general hydrophobicity) were tested and killed. The same null-gated test applied to three unrelated benchmarks (MTEB embeddings, SWE-bench code-agent evaluation, LMArena chat preference) finds a hidden axis in every one that resolves at all. Base-rate collapse: ten clinical variant-pathogenicity predictors report AUCs of 0.81-0.98 on a benchmark that is 65% pathogenic; reprojected by Bayes to the 0.1-1% prevalence a real screening population presents, the best predictor's flag precision falls from 93% to 12% (1.3% at 0.1%). The identical arithmetic, applied to AI-generated-text detection on GPT-4 output, takes the strongest detector's benchmark precision of 79% to 16% at a 5% real-world machine-text rate -- about five false accusations per genuine catch -- and shows the classifier lineage behind many commercial "AI checkers" scores AUC=0.545 on GPT-4, indistinguishable from a coin flip. Genomics and generated text share the identical failure shape because the underlying arithmetic is the same theorem; only the base rate changes. The paper releases two reusable, self-tested primitives -- a rank-dimensionality auditor and a Bayes recalibrator -- and the full analysis pipeline.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Genomics and Rare Diseases
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.