DeltaKV: Beyond Attention Magnitude for KV Cache Compression
Abstract
The growing memory footprint and computation cost of the key–value (KV) cache pose a longstanding bottleneck for efficient long-context processing. KV eviction mitigates these costs by retaining a subset of cached entries, and most existing methods decide which entries to retain based on attention scores. However, our analysis shows that attention scores exhibit both query-relevant and query-invariant components; attention over the same context exhibits substantial shared patterns across different queries. Under a limited KV budget, these shared patterns may bias KV selection toward entries that are less relevant to the current query. Motivated by this observation, we propose DeltaKV, a KV eviction method that scores cached entries by contrasting attention under the actual query with attention under a generic proxy query. This simple contrast discounts shared structure and emphasizes query-specific changes to guide KV selection, with minimal computational overhead. Experiments on LongBench-E, RULER, and MileBench show that DeltaKV degrades more gracefully than competing methods as the compression ratio increases. Across these three benchmarks, evicting 90% of KV entries costs less than 7% relative performance. Even with only 2% of the original KV cache, DeltaKV retains over 70% of full-cache performance and outperforms all evaluated baselines in every model–benchmark combination.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.