acceptodds
Under review as a conference paper at ICLR 2027

Balancing Query Coverage in KV Cache Compression

Abstract

During autoregressive decoding, the KV cache supports queries that may depend on different parts of a long context. Under a limited memory budget, selecting tokens by aggregate attention can leave some queries poorly covered even when the total retained attention is high. We study this imbalance through a distributional formulation of cache eviction, measuring query coverage by the attention probability mass assigned to retained tokens. For fixed queries and keys, restricting attention to a retained token subset is equivalent to conditioning the original attention distribution on that subset. The KL divergence from the conditioned distribution to the original distribution equals the negative logarithm of the retained probability mass, yielding an objective that maximizes the geometric mean of query coverage. Its marginal scores give greater weight to attention from queries with less retained mass. We instantiate this principle as FidKV, a lightweight online eviction policy that estimates retained attention mass from the current cache without additional training or calibration. Our analysis bounds changes in token rankings by the spread of retained mass across queries, recovering cumulative-attention rankings when coverage is uniform. Experiments on LongBench and RULER across three language-model families support this prediction. On six held-out LongBench tasks, the advantage of FidKV is most pronounced under aggressive compression: at 2% and 5% cache budgets, it outperforms the strongest evaluated baseline on most tasks, while remaining competitive at moderate budgets., respectively. Paired cache measurements further show improved coverage for the worst-served queries and reduced coverage imbalance, while deployment retains the memory and latency profile of standard KV eviction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.