acceptodds
Under review as a conference paper at ICLR 2027

Not All Queries Are Equal: Identifying Informative Queries for KV Cache Compression

Abstract

Long prompts make the KV cache the dominant memory cost, so a common remedy is to keep only part of it once prefill ends. The standard methods share one form: they sum the attention each key receives from prefill queries and keep the keys with the largest sums. They differ only in which queries they include. We show that queries differ sharply in how useful they are for this sum. A few informative queries point to the keys that generation will read, while most are noisy and do not, and the choice between them largely determines accuracy. We propose Q-Trust, which identifies informative queries from the attention of the user's request and selects the keys they point to through 2-hop attention. Without any training, Q-Trust improves accuracy over SnapKV by 1.72–2.56 points on Llama-3.1-8B, Qwen2.5-7B, and Mistral-7B-v0.3, and by 0.78–1.58 points over SnapKV with pooling. A Triton kernel that reuses the log-sum-exp from FlashAttention limits the prefill overhead to 3.5% for 16K-token prompts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.