Text Cannot Remove a Memory: Update, Retraction, and Access Control in Language Models
Language models increasingly serve as long-term memory, and systems update or retract stored facts by writing text into the context. What do these writes actually change? Across model scales (0.5B–7B) and a second architecture family, we find that, for a system that controls the model's cache, overwriting the relevant key–value state at write time acts as a replacement under our tests: behavior flips to the new value, the old value does not return when its text sources are masked, and probes do not recover it. Retraction is different: appended void notices rarely change answers in our templates; when a retraction is compiled into the record, answers flip but almost always mention the old value; and frontier models comply with the retraction yet surface the value in every compliant answer we audited. Mechanically, text reweights the continuation distribution but cannot remove a candidate, and status wording stays below the measured flip threshold. A read gate or fine-tuning restores retraction behavior, but probes recover the suppressed value and a neutral fine-tune revives it: learned retraction is behavior-level suppression, not removal. Rebuilding the context from an authorized view, so the value never enters it, passes the same audits as update; on real data, removing the superseded evidence rather than prompting changes answers. Behavior is not information: "we removed it" should meet the same evidence standard as "we changed it."
Authors
- Junchen Chen
Institutions
- Shanghai University (CN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23062085
- Primary Topic
- Topic Modeling
- Type
- preprint