Whether a KV Cache Can Be Evicted Blind Is a Property of Its Queries
Abstract
Two summary statistics of an attention head's queries, their average direction and how tightly they cluster around it, predict whether the head's KV cache can be evicted blind, that is, cut to one set of kept tokens chosen before any query arrives, for queries that continue the document. For questions appended after the document, the model family decides: under a decision rule recorded before the analysis, the prediction holds on Olmo-3 and fails on Llama, Qwen2.5 and Qwen3.5. Blind eviction works on a head only when its queries agree on which cached tokens matter, and the two statistics rank heads by that agreement, with nothing fitted, at a Spearman correlation of at least +0.89 on held-out queries in open-weight models from three families. A serving system can therefore tell before decoding which heads it can evict blind on continuation-style workloads. Measured exactly, this agreement is low and varies widely between heads of one model. A scoring rule built from the average direction adds less output error, over the best set each query could have kept, than every recent method we compare, as released, on held-out continuation queries, and on the LongBench benchmark, whose questions are appended, it is level with Expected Attention within its interval. The measurements also explain why RoCo's division of accumulated attention by the number of positions that could see each token, which kvpress already applies, cuts that excess error for H2O by 45%. They show that a guarantee met on the queries that chose the kept set breaks on many unseen ones, and they yield five practical implications for eviction. Code: https://anonymous.4open.science/r/kv-incompressibility-A706/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.