PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
Retrieval-Augmented Generation (RAG) significantly improves Large Language Models (LLMs) but introduces massive input sequences that severely bottleneck the prefill stage. While KV-cache reuse reduces redundant computation for shared document prefixes, the reusable KV working set in RAG serving can exceed GPU memory capacity, requiring KV chunks to be retained across host DRAM and SSDs. However, naive multi-tier storage extensions suffer from severe I/O bottlenecks, suboptimal eviction, and high CPU-GPU data transfer overheads, which often negate the latency benefits of cache reuse. In this paper, we propose PCR, a P refetch-enhanced C ache R euse system for low-latency RAG serving. PCR transforms passive SSD-backed storage into an active, latency-hiding memory hierarchy through three core modules: (1) a Multi-Tier Prefix-Tree Cache Manager that unifies the organization and tracking of reusable KV chunks across the entire memory hierarchy of GPU, host DRAM, and SSD; (2) a Queue-Guided Runtime Scheduler that leverages the post-retrieval waiting queue as a look-ahead signal to proactively protect hot chunks and prefetch SSD-resident data into host DRAM before execution; and (3) a Latency-Hiding KV Transfer Pipeline that overlaps fine-grained PCIe data movement and asynchronous SSD operations with model computation. Extensive evaluations across diverse models and RAG workloads show that PCR reduces TTFT over the evaluated KV-cache reuse systems in most settings, with a 2.45 × speedup over vLLM for Llama3.1-8B on A6000 under Workload 1 at 1.0 request/s, while maintaining lower tail latency under high load.
Authors
- Hengyi Zhou
- Chao Li (ORCID: https://orcid.org/0009-0006-6143-8673)
- Xinkai Wang
- Minyi Guo
- Xiaofeng Hou
- Jing Wang (ORCID: https://orcid.org/0000-0002-5513-9996)
- Peng Tang
- Wenfeng Wang
Institutions
- Guizhou University (CN)
- Shanghai Jiao Tong University (CN)
- East China Normal University (CN)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1145/3856989
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00