acceptodds
Under review as a conference paper at ICLR 2027

Fast and Quality-Preserving Non-prefix KV Cache Reuse for Large Language Model Serving

Abstract

Key-Value (KV) Cache reuse reduces prefill computation for long-context large language models (LLMs), but reusing independently encoded text chunks in non-prefix contexts degrades generation quality because cross-chunk attention is missing. Selective recomputation can mitigate this problem, but it requires both selecting effective tokens for recomputation and preventing the costs of recomputation and cache loading from offsetting the benefits of reuse. We observe that the frequency-domain structure of KV provides an effective signal for identifying tokens that help restore cross-chunk dependencies. Based on this observation, we propose CacheTune, a frequency-guided method for non-prefix KV Cache reuse. CacheTune analyzes independently encoded KV Chunks offline to generate recomputation indices shared across layers, recomputes the selected KVs in the new context, and reuses the remaining cache, without additional training or online importance analysis. The same indices also determine which KVs must be loaded, so the volume of cache transfers decreases as the recomputation ratio increases, while allowing transfer to overlap with computation. CacheTune further calibrates the recomputation ratio based on hardware computation costs and cache access costs to accommodate different storage tiers. Across mainstream models and long-context tasks, CacheTune achieves a – speedup in time-to-first-token (TTFT) and – the throughput of full recomputation while maintaining comparable generation quality. With SSD/HDD caches, the TTFT speedup still reaches –. The code is available at https://anonymous.4open.science/r/CacheTune-2123/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.