Removing the Competitor Is Not Enough: Reallocating Attention to the Current Value Restores Recall under Proactive Interference

When one key is updated again and again within a single context, language models answer with a value that has already been overwritten, long before the context window fills. Suppressing attention to the stale value reduces those answers, but no study has held the deletion fixed and varied only where the deleted attention goes, so it is unknown how much of a repair comes from the deletion and how much from the destination. Here we delete the answer position’s attention to a key’s superseded values by one rule at the same positions under every condition and send it to a different destination in each, in five models from four families (1.5B–8.79B parameters). Deleting it without a destination never restores clean accuracy: it recovers from 0.25 of the accuracy gap in Qwen2.5-1.5B to 0.88 in Granite-4.2-8B. Sending it to the queried key’s current value restores clean accuracy in every model and beats spreading it elsewhere, by +0.63 [+0.54, +0.72] in Qwen2.5-1.5B down to +0.18 [+0.11, +0.26] in Granite-4.2-8B (95% intervals). Sent instead to another key’s current value, it restores nothing and the model answers that value; sent to another key’s earlier value or to a key word, it answers that token. Without any deletion, swapping in the attention from a matched prompt in which the same old values are assigned to a new key raises accuracy more than swapping in that prompt’s values, by +0.558 to +0.835: the failure is in where the answer position reads. Removing the competitor is not enough: in these models, only sending the deleted attention to the current value restores clean recall.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23153676
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Removing the Competitor Is Not Enough: Reallocating Attention to the Current Value Restores Recall under Proactive Interference

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

Removing the Competitor Is Not Enough: Reallocating Attention to the Current Value Restores Recall under Proactive Interference

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

When one key is updated again and again within a single context, language models answer with a value that has already been overwritten, long before the context window fills. Suppressing attention to the stale value reduces those answers, but no study has held the deletion fixed and varied only where the deleted attention goes, so it is unknown how much of a repair comes from the deletion and how much from the destination. Here we delete the answer position’s attention to a key’s superseded values by one rule at the same positions under every condition and send it to a different destination in each, in five models from four families (1.5B–8.79B parameters). Deleting it without a destination never restores clean accuracy: it recovers from 0.25 of the accuracy gap in Qwen2.5-1.5B to 0.88 in Granite-4.2-8B. Sending it to the queried key’s current value restores clean accuracy in every model and beats spreading it elsewhere, by +0.63 [+0.54, +0.72] in Qwen2.5-1.5B down to +0.18 [+0.11, +0.26] in Granite-4.2-8B (95% intervals). Sent instead to another key’s current value, it restores nothing and the model answers that value; sent to another key’s earlier value or to a key word, it answers that token. Without any deletion, swapping in the attention from a matched prompt in which the same old values are assigned to a new key raises accuracy more than swapping in that prompt’s values, by +0.558 to +0.835: the failure is in where the answer position reads. Removing the competitor is not enough: in these models, only sending the deleted attention to the current value restores clean recall.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW)
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Removing the Competitor Is Not Enough: Reallocating Attention to the Current Value Restores Recall under Proactive Interference — Ya-Fen Yeh, Guan-Yuan Chen · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS