Key-value-free transformers: Training cacheless decoders via residual stream stabilisation

Serving large language models is memory-bound: the key-value (KV) cache grows linearly with sequence length. Cached keys and values are deterministic projections of the residual stream, so the cache carries nothing the residual lacks: the KV-Free identity. Residual checkpoints reconstruct them losslessly, exposing a memory-for-compute frontier. The identity holds exactly on sixteen models, eleven families, 124 million to 8 billion parameters. One residual per token compresses stored state 5.5 to 64 times at 100% token match; five lossy eviction baselines fall to 5 to 28%. Against other lossless schedules it is the maximum-compression interior point A cacheless decoder holds lower peak memory on three 8-billion-parameter models and serves 32,768 tokens on one where a resident cache runs out. Recomputing the whole prefix, the costliest schedule, bounds the price at 114 times the energy per token; recomputing a quarter of the layers cuts that 3.7-fold so the saving suits batched long-context serving, not batch one. In a 47-million-parameter model under a ten-minute budget, raising the adaptive moment estimation optimiser’s second-moment decay from 0.95 to 0.9999 lowers validation bits-per-byte by 10.3%; we report the effective-rank drop from 84 to 77 as a correlate. Cache elimination becomes an operating point, priced in recomputation.

Authors

Institutions

Publication Details

Journal
Engineering Applications of Artificial Intelligence
Published
2026-10-07
DOI
https://doi.org/10.1016/j.engappai.2026.116397
Primary Topic
Natural Language Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Key-value-free transformers: Training cacheless decoders via residual stream stabilisation

Jiashu Zhang, Razan Alharith, Kaleem Ullah Qasim, Muhammad Kafeel Shaheen
Engineering Applications of Artificial Intelligence
Natural Language Processing Techniques
article

Key-value-free transformers: Training cacheless decoders via residual stream stabilisation

Jiashu Zhang, Razan Alharith, Kaleem Ullah Qasim, Muhammad Kafeel Shaheen
article en

Abstract

Serving large language models is memory-bound: the key-value (KV) cache grows linearly with sequence length. Cached keys and values are deterministic projections of the residual stream, so the cache carries nothing the residual lacks: the KV-Free identity. Residual checkpoints reconstruct them losslessly, exposing a memory-for-compute frontier. The identity holds exactly on sixteen models, eleven families, 124 million to 8 billion parameters. One residual per token compresses stored state 5.5 to 64 times at 100% token match; five lossy eviction baselines fall to 5 to 28%. Against other lossless schedules it is the maximum-compression interior point A cacheless decoder holds lower peak memory on three 8-billion-parameter models and serves 32,768 tokens on one where a resident cache runs out. Recomputing the whole prefix, the costliest schedule, bounds the price at 114 times the energy per token; recomputing a quarter of the layers cuts that 3.7-fold so the saving suits batched long-context serving, not batch one. In a 47-million-parameter model under a ten-minute budget, raising the adaptive moment estimation optimiser’s second-moment decay from 0.95 to 0.9999 lowers validation bits-per-byte by 10.3%; we report the effective-rank drop from 84 to 77 as a correlate. Cache elimination becomes an operating point, priced in recomputation.

Engineering Applications of Artificial IntelligenceVol. 185
Southwest Jiaotong University (CN)
Openalex Percentile: Top 12%
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Key-value-free transformers: Training cacheless decoders via residual stream stabilisation — Jiashu Zhang, Razan Alharith, et al. · Engineering Applications of Artificial Intelligence (2026) | TGRS Research Map | TGRS