acceptodds
Under review as a conference paper at ICLR 2027

ER-KV: Eviction Risk-Aware KV Cache Compression for Efficient Long Reasoning

Abstract

Reasoning language models generate long chain-of-thought sequences, causing KV-cache memory to grow substantially during decoding. KV-cache eviction mitigates this cost, but deciding which tokens to retain requires accounting for both future queries and the attention-output error caused by token removal. Existing methods address either future attention allocation or eviction-induced attention-output error, but do not jointly account for both when making eviction decisions. We propose Eviction Risk-Aware KV Cache Compression (ER-KV), which ranks cached tokens by their expected eviction-induced attention-output error over future queries. ER-KV models the future-query distribution with a Gaussian fitted to observed pre-RoPE query statistics. To estimate eviction risk, which depends nonlinearly on each query, it samples virtual queries to represent this distribution, applies RoPE at multiple future positions, and aggregates the resulting eviction errors. Cholesky factorization reduces sampling overhead while yielding the same Gaussian sampling distribution as eigendecomposition. Under repeated pruning, ER-KV reduces attention-output MSE by 24.4% and output-distribution KL divergence by 48.6% relative to TriAttention. Across AIME24, AIME25, and MATH-500 with three reasoning models, ER-KV matches or outperforms TriAttention in eight of nine model-benchmark settings, improving Qwen3-8B accuracy on AIME24 from 42.1% to 49.2%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.