Attention Has a Coordinate but No Commitment: The Selection Gap in Softmax Attention, with Two Measurements on GPT-2

{"Summary.":[0],"Softmax":[1],"attention":[2,92,346],"is":[3,49,55,81,97,185,255,268,307],"often":[4],"described":[5],"as":[6,22,403],"\\"smooth\\":":[7],"every":[8],"unmasked":[9],"key":[10],"receives":[11],"nonzero":[12],"weight.":[13],"An":[14],"earlier":[15,187,331],"paper":[16,37,324,337],"by":[17,51,217,234,309,319],"this":[18],"author":[19],"took":[20],"that":[21,291,302,336,372],"the":[23,47,59,67,103,113,118,159,163,167,186,212,225,228,240,269,284,305,323,330,333,354,370,392,413,415,417,420,425],"structural":[24],"origin":[25],"of":[26,62,105,121,125,132,151,166,198,215,279,358],"hallucination":[27,352],"and":[28,46,75,193,201,204,222,242,246,276,314,343,350,369,400,424],"proposed":[29],"a":[30,63,71,76,82,86,95,98,106,129,149,173,179,207,263,292,359,366],"learned":[31],"hard":[32,180],"threshold":[33,96],"per":[34],"head.":[35],"This":[36],"supersedes":[38],"it.":[39,375],"The":[40,79,91,145,288,326,410],"diagnosis":[41],"was":[42,386,407],"wrong":[43],"in":[44,123,137,142,162,365],"mechanism":[45],"remedy":[48],"contradicted":[50],"measurement;":[52],"what":[53],"survives":[54],"narrower:":[56],"softmax":[57],"supplies":[58],"shared":[60],"coordinate":[61],"selection":[64,83,363],"but":[65],"not":[66,85,256,318],"two":[68,426],"other":[69],"parts,":[70],"commitment":[72],"with":[73,224,247,362,381],"memory":[74,316],"persistent":[77],"label.":[78],"gap":[80],"gap,":[84],"smoothness":[87],"gap.":[88],"Measurement":[89,170],"M1.":[90],"mass":[93,230],"below":[94],"shared,":[99],"common-mode":[100,114],"quantity":[101],"across":[102],"heads":[104],"layer":[107,138,143],"rather":[108],"than":[109,298],"per-head":[110,181],"structure.":[111],"Position-detrended,":[112],"variance":[115],"fraction":[116],"exceeds":[117],"registered":[119,147,289,397],"bar":[120],"0.2":[122],"10":[124],"12":[126],"layers":[127],"against":[128],"shuffle":[130],"null":[131],"0.083,":[133],"rising":[134],"from":[135,262,329],"0.09":[136],"0":[139],"to":[140],"0.48":[141],"11.":[144],"stronger":[146],"form,":[148],"median":[150],"at":[152,156,196,236,239,250,283,301],"least":[153],"0.3,":[154],"fails":[155],"0.278,":[157],"so":[158,253,304],"reading":[160],"holds":[161],"deep":[164],"half":[165],"network":[168],"only.":[169],"M2.":[171],"At":[172],"matched":[174],"retained-key":[175],"budget,":[176,252,303],"closed":[177],"loop,":[178],"level":[182,293],"threshold,":[183],"which":[184,267],"paper's":[188],"proposal,":[189],"costs":[190,219,273],"38.4,":[191],"4.0":[192],"0.1":[194],"perplexity":[195,275],"budgets":[197,245],"10,":[199],"20":[200,241,285],"40":[202,243],"percent":[203,244,286],"never":[205],"improves":[206],"named-entity":[208],"continuation":[209],"proxy.":[210],"Selecting":[211],"same":[213],"number":[214],"keys":[216],"rank":[218,235],"0.5,":[220],"-0.1":[221,223],"proxy":[226],"unchanged:":[227],"low-attention":[229],"can":[231],"be":[232],"removed":[233],"no":[237,248,405],"cost":[238,296],"benefit":[249],"any":[251],"it":[254],"where":[257],"misplaced":[258],"content":[259],"lives.":[260],"Eviction":[261],"leaky,":[264],"read-refreshed":[265],"store,":[266],"constructive":[270],"proposal's":[271],"memory,":[272],"7.9":[274],"five":[277],"points":[278],"in-context":[280],"name":[281],"recall":[282],"budget.":[287],"prediction":[290],"gate":[294],"would":[295,373],"more":[297],"eviction":[299],"failed":[300,399],"store":[306],"justified":[308],"its":[310,315],"persistent,":[311],"erasable":[312],"label":[313],"footprint,":[317],"perplexity.":[320],"What":[321],"else":[322],"states.":[325],"claims":[327],"withdrawn":[328],"paper;":[332],"prior":[334],"art":[335],"omitted,":[338],"namely":[339],"sparsemax,":[340],"entmax,":[341],"rectified":[342],"top-k":[344],"attention,":[345],"sinks,":[347],"shared-workspace":[348],"transformers":[349],"calibration-based":[351],"bounds;":[353],"quantitative":[355],"building":[356],"blocks":[357],"store-plus-rank-selector":[360],"architecture":[361],"run":[364],"separate":[367],"phase;":[368],"measurement":[371,418,427],"falsify":[374],"Method.":[376],"Both":[377],"measurements":[378],"were":[379],"pre-registered":[380],"falsification":[382],"criteria":[383,398],"before":[384],"either":[385],"run,":[387],"on":[388],"GPT-2":[389],"small":[390],"over":[391],"wikitext-2-raw-v1":[393],"validation":[394],"split.":[395],"Two":[396],"are":[401],"reported":[402],"failures;":[404],"criterion":[406],"adjusted":[408],"afterwards.":[409],"deposit":[411],"contains":[412],"paper,":[414],"pre-registration,":[416],"code,":[419],"raw":[421],"result":[422],"files":[423],"reports.":[428]}

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-17
DOI
https://doi.org/10.5281/zenodo.22820231
Primary Topic
Computability, Logic, AI Algorithms
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Attention Has a Coordinate but No Commitment: The Selection Gap in Softmax Attention, with Two Measurements on GPT-2

Michael Bieg
Zenodo (CERN European Organization for Nuclear Research)
Computability, Logic, AI Algorithms
preprint

Attention Has a Coordinate but No Commitment: The Selection Gap in Softmax Attention, with Two Measurements on GPT-2

Michael Bieg
preprint en

Abstract

Summary. Softmax attention is often described as "smooth": every unmasked key receives nonzero weight. An earlier paper by this author took that as the structural origin of hallucination and proposed a learned hard threshold per head. This paper supersedes it. The diagnosis was wrong in mechanism and the remedy is contradicted by measurement; what survives is narrower: softmax supplies the shared coordinate of a selection but not the two other parts, a commitment with memory and a persistent label. The gap is a selection gap, not a smoothness gap. Measurement M1. The attention mass below a threshold is a shared, common-mode quantity across the heads of a layer rather than per-head structure. Position-detrended, the common-mode variance fraction exceeds the registered bar of 0.2 in 10 of 12 layers against a shuffle null of 0.083, rising from 0.09 in layer 0 to 0.48 in layer 11. The stronger registered form, a median of at least 0.3, fails at 0.278, so the reading holds in the deep half of the network only. Measurement M2. At a matched retained-key budget, closed loop, a hard per-head level threshold, which is the earlier paper's proposal, costs 38.4, 4.0 and 0.1 perplexity at budgets of 10, 20 and 40 percent and never improves a named-entity continuation proxy. Selecting the same number of keys by rank costs 0.5, -0.1 and -0.1 with the proxy unchanged: the low-attention mass can be removed by rank at no cost at the 20 and 40 percent budgets and with no benefit at any budget, so it is not where misplaced content lives. Eviction from a leaky, read-refreshed store, which is the constructive proposal's memory, costs 7.9 perplexity and five points of in-context name recall at the 20 percent budget. The registered prediction that a level gate would cost more than eviction failed at that budget, so the store is justified by its persistent, erasable label and its memory footprint, not by perplexity. What else the paper states. The claims withdrawn from the earlier paper; the prior art that paper omitted, namely sparsemax, entmax, rectified and top-k attention, attention sinks, shared-workspace transformers and calibration-based hallucination bounds; the quantitative building blocks of a store-plus-rank-selector architecture with selection run in a separate phase; and the measurement that would falsify it. Method. Both measurements were pre-registered with falsification criteria before either was run, on GPT-2 small over the wikitext-2-raw-v1 validation split. Two registered criteria failed and are reported as failures; no criterion was adjusted afterwards. The deposit contains the paper, the pre-registration, the measurement code, the raw result files and the two measurement reports.

Zenodo (CERN European Organization for Nuclear Research)
Partnerships for the goals
Computability, Logic, AI Algorithms
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.