Fast and Quality-Preserving Non-prefix KV Cache Reuse for Large Language Model Serving
Abstract
Key-Value (KV) Cache reuse reduces prefill computation for long-context large language models (LLMs), but reusing independently encoded text chunks in non-prefix contexts degrades generation quality because cross-chunk attention is missing. Selective recomputation can mitigate this problem, but it requires both selecting effective tokens for recomputation and preventing the costs of recomputation and cache loading from offsetting the benefits of reuse. We observe that the frequency-domain structure of KV provides an effective signal for identifying tokens that help restore cross-chunk dependencies. Based on this observation, we propose CacheTune, a frequency-guided method for non-prefix KV Cache reuse. CacheTune analyzes independently encoded KV Chunks offline to generate recomputation indices shared across layers, recomputes the selected KVs in the new context, and reuses the remaining cache, without additional training or online importance analysis. The same indices also determine which KVs must be loaded, so the volume of cache transfers decreases as the recomputation ratio increases, while allowing transfer to overlap with computation. CacheTune further calibrates the recomputation ratio based on hardware computation costs and cache access costs to accommodate different storage tiers. Across mainstream models and long-context tasks, CacheTune achieves a – speedup in time-to-first-token (TTFT) and – the throughput of full recomputation while maintaining comparable generation quality. With SSD/HDD caches, the TTFT speedup still reaches –. The code is available at https://anonymous.4open.science/r/CacheTune-2123/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.