acceptodds
Under review as a conference paper at ICLR 2027

Query-Aware KV Block Selection from Cache Landmarks

Abstract

Autoregressive language models repeatedly read a growing key-value (KV) cache when generating from long contexts. Query-aware block selection retains the cache but limits attention to blocks chosen for each new token; its quality depends on ranking those blocks well. We introduce KV-Selector, a small learned scorer for frozen language models that reuses the minimum and maximum key summaries stored by Quest, an earlier selection method. It trains from the model's own attention in minutes, uses fewer than 10,000 parameters per layer and adds no summaries beyond Quest's. Across eleven models, it better preserves next-token predictions than the compared training-free rules in 78 of 82 settings that restrict one layer at a time. On four models with a 640-token read budget, it lowers negative log-likelihood on PG19, a benchmark of language modelling over books, by 0.013 to 0.039 nats relative to the best training-free rule making one block choice per layer. Retrieval gains are less consistent. With the cache in host memory, selecting blocks accelerates decoding relative to synchronously copying the whole cache, and on PG19 at 131k tokens KV-Selector keeps a lower negative log-likelihood than the compared training-free rules. Our unfused implementation remains slower than full attention with the cache on the GPU.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.