acceptodds
Under review as a conference paper at ICLR 2027

QueEn: Query Enrichment for KV Cache Eviction in Long Reasoning

Abstract

Key-value (KV) cache eviction alleviates the memory bottleneck of long-reasoning inference by selectively retaining important KV pairs. Most eviction methods score each cached KV pair by the attention it receives from the queries of the most recent tokens. However, these scores can undervalue KV pairs that future decoding still needs, since the queries cover only recent context and, under causal masking, each cannot see later tokens in the window. To address these limitations, we propose Query Enrichment (QUEEN), a training-free method that refines the queries of the most recent tokens so that each reflects all of them (intra-window refinement), and reuses refined queries from earlier reasoning and the input prompt (cross-context reuse), with additional batch-scalable reuse scheduling, which reduces the peak GPU memory of the reused queries for higher-throughput serving. On three long- reasoning benchmarks, the proposed QUEEN consistently outperforms existing eviction methods. Especially on LiveCodeBench, it matches full-KV accuracy with 5.2× less KV memory and 4.1× higher throughput.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.