QueEn: Query Enrichment for KV Cache Eviction in Long Reasoning
Abstract
Key-value (KV) cache eviction alleviates the memory bottleneck of long-reasoning inference by selectively retaining important KV pairs. Most eviction methods score each cached KV pair by the attention it receives from the queries of the most recent tokens. However, these scores can undervalue KV pairs that future decoding still needs, since the queries cover only recent context and, under causal masking, each cannot see later tokens in the window. To address these limitations, we propose Query Enrichment (QUEEN), a training-free method that refines the queries of the most recent tokens so that each reflects all of them (intra-window refinement), and reuses refined queries from earlier reasoning and the input prompt (cross-context reuse), with additional batch-scalable reuse scheduling, which reduces the peak GPU memory of the reused queries for higher-throughput serving. On three long- reasoning benchmarks, the proposed QUEEN consistently outperforms existing eviction methods. Especially on LiveCodeBench, it matches full-KV accuracy with 5.2× less KV memory and 4.1× higher throughput.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.