Key-value-free transformers: Training cacheless decoders via residual stream stabilisation
Serving large language models is memory-bound: the key-value (KV) cache grows linearly with sequence length. Cached keys and values are deterministic projections of the residual stream, so the cache carries nothing the residual lacks: the KV-Free identity. Residual checkpoints reconstruct them losslessly, exposing a memory-for-compute frontier. The identity holds exactly on sixteen models, eleven families, 124 million to 8 billion parameters. One residual per token compresses stored state 5.5 to 64 times at 100% token match; five lossy eviction baselines fall to 5 to 28%. Against other lossless schedules it is the maximum-compression interior point A cacheless decoder holds lower peak memory on three 8-billion-parameter models and serves 32,768 tokens on one where a resident cache runs out. Recomputing the whole prefix, the costliest schedule, bounds the price at 114 times the energy per token; recomputing a quarter of the layers cuts that 3.7-fold so the saving suits batched long-context serving, not batch one. In a 47-million-parameter model under a ten-minute budget, raising the adaptive moment estimation optimiser’s second-moment decay from 0.95 to 0.9999 lowers validation bits-per-byte by 10.3%; we report the effective-rank drop from 84 to 77 as a correlate. Cache elimination becomes an operating point, priced in recomputation.
Authors
- Jiashu Zhang (ORCID: https://orcid.org/0000-0003-2086-0991)
- Razan Alharith (ORCID: https://orcid.org/0000-0003-4296-3510)
- Kaleem Ullah Qasim (ORCID: https://orcid.org/0000-0002-0102-3816)
- Muhammad Kafeel Shaheen (ORCID: https://orcid.org/0009-0009-0619-1737)
Institutions
- Southwest Jiaotong University (CN)
Publication Details
- Journal
- Engineering Applications of Artificial Intelligence
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1016/j.engappai.2026.116397
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00