Keys Decide Attention, Values Decide Contribution: Value-Gated KV Cache Eviction at 20× Compression
Abstract
Long documents are queried many times, so serving systems keep a document's key–value (KV) cache and compress it once, before any question arrives. Existing query-agnostic eviction methods rank cache entries by their keys alone and collapse under severe compression: with 5% of the cache retained, the best of them scores 34 on LongBench passage retrieval, against 100 for the full cache. Key-only scores ignore half of attention. Keys decide whether an entry is attended; values decide what an attended entry contributes, and a distinctive key with a small projected value contributes almost nothing. We propose VG-Evict (Value-Gated Eviction), a training-free rule that multiplies KeyDiff's normalized key-geometry score by the normalized magnitude of each entry's output-projected value, with a small floor and a positional reserve, at cost linear in context length. At 5%/7.5% retention, VG-Evict scores 50.5/70.5 on Qwen2.5-7B, 16.5/11.5 points above Compactor and 7.5/7.0 above CriticalKV applied to KeyDiff, and gains 14.0/17.0 points over Compactor on Mistral-7B; every paired bootstrap interval excludes zero. Controlled ablations attribute the gain to token-aligned value information: gating adds 18.0/23.5 points over KeyDiff at a matched reserve, shuffling the gates among tokens erases the gain, and a single ranking stage matches the best of CriticalKV's two-stage configurations without a stage-ratio hyperparameter. One compressed cache also serves a whole session: on 100 RULER documents at 10% retention it answers 97.5% of 400 independent questions with a tenth of the KV memory. The gains concentrate on passage identification and sparse lookup; for free-form answer extraction, leverage-score selection remains preferable. Code: https://anonymous.4open.science/status/VG-Evict-Anonymous-Code-706D.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.