KVgrad: Query-Agnostic KV Cache Eviction via Gradient-based Global Importance Scoring
Abstract
Key-Value (KV) cache eviction mitigates the memory overhead of Large Language Models (LLMs) by discarding cached KV pairs that are less relevant to future token generation. However, existing eviction methods typically rely on local heuristics, scoring cached entries by their immediate contribution to the attention output of the same layer, while overlooking their downstream impact across subsequent layers. To address this, we propose KVgrad, a query-agnostic eviction method that scores each cache entry by its global impact on the model's final outputs. Based on a simple analysis using a first-order Taylor approximation and the chain rule, KVgrad factorizes the impact of cached value into a linear local component and a gradient-based downstream sensitivity term. Then, we apply a magnitude-aware scoring scheme to stabilize the importance score. Evaluations on RULER and LongBench demonstrate that KVgrad achieves an average maximum compression ratio of 5.9x with negligible performance loss (<2%), significantly outperforming strong local-only state-of-the-art baselines. Furthermore, we demonstrate that our signal is highly amenable to distillation; by training a lightweight MLP to predict KVgrad scores from local hidden-state features, we enable high-fidelity, on-the-fly compression at negligible computational cost. Finally, we demonstrate the transferability of KVgrad to vision-language models; evaluations on image and long-context video benchmarks suggest a unified path for importance-aware KV cache compression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.