Not All Queries Are Equal: Identifying Informative Queries for KV Cache Compression
Abstract
Long prompts make the KV cache the dominant memory cost, so a common remedy is to keep only part of it once prefill ends. The standard methods share one form: they sum the attention each key receives from prefill queries and keep the keys with the largest sums. They differ only in which queries they include. We show that queries differ sharply in how useful they are for this sum. A few informative queries point to the keys that generation will read, while most are noisy and do not, and the choice between them largely determines accuracy. We propose Q-Trust, which identifies informative queries from the attention of the user's request and selects the keys they point to through 2-hop attention. Without any training, Q-Trust improves accuracy over SnapKV by 1.72–2.56 points on Llama-3.1-8B, Qwen2.5-7B, and Mistral-7B-v0.3, and by 0.78–1.58 points over SnapKV with pooling. A Triton kernel that reuses the log-sum-exp from FlashAttention limits the prefill overhead to 3.5% for 16K-token prompts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.