AdapTQ-GPU: GPU-Resident Adaptive KV-Cache Compression for Memory-Bounded Long-Context Inference
AdapTQ-GPU is a GPU-resident extension of the AdapTQ approach for adaptive key-value (KV) cache compression in long-context large language model inference. The system combines GPU-resident compressed KV storage with memory-bounded streaming prefill and Triton-based compressed-cache decoding. The compression pipeline applies Rademacher sign randomization, a Fast Walsh-Hadamard Transform (FWHT), asymmetric scalar quantization, and packed KV storage while retaining Layer 0 in uncompressed FP16. The implementation was evaluated on an NVIDIA GeForce RTX 5050 Laptop GPU with approximately 8 GB of GPU memory using Qwen2-1.5B-Instruct. Experiments evaluated contexts from 16K through 131K tokens using chunked prefill. At 131K tokens, dense FP16 KV storage required approximately 3.58 GB for the persistent KV cache, while the implemented AdapTQ configuration required approximately 830 MB including the uncompressed Layer-0 cache, corresponding to approximately 4.3× total KV-cache compression. The compressed configuration completed the 131K-token experiment using approximately 4.76 GB of observed physical GPU memory, while dense FP16 failed under the tested hardware and software configuration. The results demonstrate a memory-capacity advantage for GPU-resident KV-cache compression under memory-bounded long-context inference. The work does not claim lossless equivalence or inference-speed improvement. Numerical divergence increases at longer contexts, with a final-token logit cosine similarity of 0.9067 at 64K tokens in the reported comparison. This work is a research prototype and is not presented as a production-ready vLLM or CUDA implementation. Source code and experimental materials:https://github.com/l3tchupkt/adaptq
Authors
- K Lakshmikanthan (ORCID: https://orcid.org/0009-0003-2087-6229)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23233683
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- preprint