RECAP-KV: Single-Prefill Approximation of Reconstruction Importance for Reusable KV Cache Eviction
Abstract
Key–value (KV) caching avoids recomputing the attention keys and values of processed tokens during generation. It also lets a long document, once processed, serve many subsequent queries. However, the cache grows with context length, limiting batch size and the number of caches that can remain resident in GPU memory. KV eviction addresses this by retaining only a subset of the cached KV pairs. Yet identifying KV pairs that preserve answer quality for unknown future queries is challenging, and reliable importance estimation is expensive for long contexts. We present RECAP-KV, a query-agnostic eviction method that reduces this cost while maintaining competitive answer quality. It applies the model's existing Q/K projections and normalization to attention-input states captured during prefill and uses rotary positional embeddings at virtual positions to approximate the attention that cached keys would receive during context reconstruction, without an additional transformer pass or a learned importance predictor. RECAP-KV allocates the cache budget across layers, heads, and context blocks, and its Lookback Reranking refines token selection without changing the number of KV pairs retained in each block. On 12 tasks and 6 models from the Qwen, Llama, and GPT-OSS families, RECAP-KV achieves the best or near-best mean task scores among query-agnostic methods at 15% and 30% KV retention, while its prefill-and-scoring time at 128K tokens is 37–46% lower than that of the strongest baseline under matched settings on B300 GPUs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.