Title RER-Q: Reading a Quantized KV Cache with Fresh Samples at Every Decoding Step

In long-context decoding the model reads its whole key–value (KV) cache at every step, and the bytes read per step set the step time, most of all when the cache is offloaded to host memory and read over PCIe. We present RER-Q (Resample-Every-Read with Quantized storage). It stores the far part of the KV cache at high precision (6–8 bits), keeps a 2-bit sketch of the keys on the GPU, and at every decoding step reads a freshly drawn subset of the tokens. A token is drawn with probability proportional to its attention mass as estimated from the sketch, and the read is corrected by inverse-probability weights. The design follows from the kind of error each choice produces. In sampled decoding the token of a step is drawn from the mixture of next-token distributions over that step's read randomness. Error that is redrawn at every step enters this mixture at second order only; fixed error, such as storage quantization or a deterministic selection, enters at first order. Our measurements indicate that, for the same read budget, it is better to read fewer tokens, at higher precision, and to draw them anew each time. A reader whose random numbers are fixed once, for all requests, is a single realization; its KL is about 12 times the per-token KL that the mixture of fresh draws adds along the sequence. At equal read budget on short contexts, the mixture KL of RER-Q is 0.36–0.90 times that of vAttention and up to two orders of magnitude (20–164 times) below that of KIVI-style dense quantization. With the cache offloaded (Qwen3-1.7B, 16k and 32k tokens, all methods reimplemented in one engine), RER-Q is 20–25% faster than vAttention at a KL ratio of 0.88 (95% interval 0.68–1.12, with the tasks as the unit). Its KL is about ten times lower than that of deterministic top-k at about the same time per token and, at 15% more storage, lower than that of 8-bit KIVI while it reads 40% of the bits and runs 1.7–1.9 times faster. At the read budget of 2-bit quantization it is 4.4–5.3 times faster than reading the fp16 cache, at a mixture KL of 1.0–1.5×10⁻⁴. On LongBench and RULER with Llama-3.1-8B (five items per task), RER-Q keeps the fp16 score where 2-bit KIVI at the same read budget loses 6 points on RULER. Task scores do not separate RER-Q from vAttention, but RER-Q reproduces the fp16 answer more often (150 against 143 of 160 items). RER-Q keeps the whole cache in host memory, at about half the size of fp16.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23115647
Primary Topic
Parallel Computing and Optimization Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Title RER-Q: Reading a Quantized KV Cache with Fresh Samples at Every Decoding Step

Yehoon Choi
Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
preprint

Title RER-Q: Reading a Quantized KV Cache with Fresh Samples at Every Decoding Step

Yehoon Choi
preprint en

Abstract

In long-context decoding the model reads its whole key–value (KV) cache at every step, and the bytes read per step set the step time, most of all when the cache is offloaded to host memory and read over PCIe. We present RER-Q (Resample-Every-Read with Quantized storage). It stores the far part of the KV cache at high precision (6–8 bits), keeps a 2-bit sketch of the keys on the GPU, and at every decoding step reads a freshly drawn subset of the tokens. A token is drawn with probability proportional to its attention mass as estimated from the sketch, and the read is corrected by inverse-probability weights. The design follows from the kind of error each choice produces. In sampled decoding the token of a step is drawn from the mixture of next-token distributions over that step's read randomness. Error that is redrawn at every step enters this mixture at second order only; fixed error, such as storage quantization or a deterministic selection, enters at first order. Our measurements indicate that, for the same read budget, it is better to read fewer tokens, at higher precision, and to draw them anew each time. A reader whose random numbers are fixed once, for all requests, is a single realization; its KL is about 12 times the per-token KL that the mixture of fresh draws adds along the sequence. At equal read budget on short contexts, the mixture KL of RER-Q is 0.36–0.90 times that of vAttention and up to two orders of magnitude (20–164 times) below that of KIVI-style dense quantization. With the cache offloaded (Qwen3-1.7B, 16k and 32k tokens, all methods reimplemented in one engine), RER-Q is 20–25% faster than vAttention at a KL ratio of 0.88 (95% interval 0.68–1.12, with the tasks as the unit). Its KL is about ten times lower than that of deterministic top-k at about the same time per token and, at 15% more storage, lower than that of 8-bit KIVI while it reads 40% of the bits and runs 1.7–1.9 times faster. At the read budget of 2-bit quantization it is 4.4–5.3 times faster than reading the fp16 cache, at a mixture KL of 1.0–1.5×10⁻⁴. On LongBench and RULER with Llama-3.1-8B (five items per task), RER-Q keeps the fp16 score where 2-bit KIVI at the same read budget loses 6 points on RULER. Task scores do not separate RER-Q from vAttention, but RER-Q reproduces the fp16 answer more often (150 against 143 of 160 items). RER-Q keeps the whole cache in host memory, at about half the size of fp16.

Zenodo (CERN European Organization for Nuclear Research)
Kyung Hee University (KR)
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.