AdapTQ-GPU: GPU-Resident Adaptive KV-Cache Compression for Memory-Bounded Long-Context Inference

AdapTQ-GPU is a GPU-resident extension of the AdapTQ approach for adaptive key-value (KV) cache compression in long-context large language model inference. The system combines GPU-resident compressed KV storage with memory-bounded streaming prefill and Triton-based compressed-cache decoding. The compression pipeline applies Rademacher sign randomization, a Fast Walsh-Hadamard Transform (FWHT), asymmetric scalar quantization, and packed KV storage while retaining Layer 0 in uncompressed FP16. The implementation was evaluated on an NVIDIA GeForce RTX 5050 Laptop GPU with approximately 8 GB of GPU memory using Qwen2-1.5B-Instruct. Experiments evaluated contexts from 16K through 131K tokens using chunked prefill. At 131K tokens, dense FP16 KV storage required approximately 3.58 GB for the persistent KV cache, while the implemented AdapTQ configuration required approximately 830 MB including the uncompressed Layer-0 cache, corresponding to approximately 4.3× total KV-cache compression. The compressed configuration completed the 131K-token experiment using approximately 4.76 GB of observed physical GPU memory, while dense FP16 failed under the tested hardware and software configuration. The results demonstrate a memory-capacity advantage for GPU-resident KV-cache compression under memory-bounded long-context inference. The work does not claim lossless equivalence or inference-speed improvement. Numerical divergence increases at longer contexts, with a final-token logit cosine similarity of 0.9067 at 64K tokens in the reported comparison. This work is a research prototype and is not presented as a production-ready vLLM or CUDA implementation. Source code and experimental materials:https://github.com/l3tchupkt/adaptq

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23233683
Primary Topic
Parallel Computing and Optimization Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

AdapTQ-GPU: GPU-Resident Adaptive KV-Cache Compression for Memory-Bounded Long-Context Inference

K Lakshmikanthan
Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
preprint

AdapTQ-GPU: GPU-Resident Adaptive KV-Cache Compression for Memory-Bounded Long-Context Inference

K Lakshmikanthan
preprint en

Abstract

AdapTQ-GPU is a GPU-resident extension of the AdapTQ approach for adaptive key-value (KV) cache compression in long-context large language model inference. The system combines GPU-resident compressed KV storage with memory-bounded streaming prefill and Triton-based compressed-cache decoding. The compression pipeline applies Rademacher sign randomization, a Fast Walsh-Hadamard Transform (FWHT), asymmetric scalar quantization, and packed KV storage while retaining Layer 0 in uncompressed FP16. The implementation was evaluated on an NVIDIA GeForce RTX 5050 Laptop GPU with approximately 8 GB of GPU memory using Qwen2-1.5B-Instruct. Experiments evaluated contexts from 16K through 131K tokens using chunked prefill. At 131K tokens, dense FP16 KV storage required approximately 3.58 GB for the persistent KV cache, while the implemented AdapTQ configuration required approximately 830 MB including the uncompressed Layer-0 cache, corresponding to approximately 4.3× total KV-cache compression. The compressed configuration completed the 131K-token experiment using approximately 4.76 GB of observed physical GPU memory, while dense FP16 failed under the tested hardware and software configuration. The results demonstrate a memory-capacity advantage for GPU-resident KV-cache compression under memory-bounded long-context inference. The work does not claim lossless equivalence or inference-speed improvement. Numerical divergence increases at longer contexts, with a final-token logit cosine similarity of 0.9067 at 64K tokens in the reported comparison. This work is a research prototype and is not presented as a production-ready vLLM or CUDA implementation. Source code and experimental materials:https://github.com/l3tchupkt/adaptq

Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

AdapTQ-GPU: GPU-Resident Adaptive KV-Cache Compression for Memory-Bounded Long-Context Inference — K Lakshmikanthan · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS