acceptodds
Under review as a conference paper at ICLR 2027

KVScout: Efficient Long-Context Decoding with Sentence-Level KV Retrieval

Abstract

Long-context models show strong capabilities in processing long inputs but incur substantial memory and computational costs due to the growing KV cache. KV retrieval methods reduce these costs by keeping the full context KV cache in CPU memory and retrieving only a small subset to the GPU for attention. However, frequent KV selection and transfer can offset sparse-attention savings and limit decoding speedup, motivating fewer retrievals during decoding. Yet, existing selectors are designed for next-token retrieval, and reusing their selections across more decoding steps degrades generation quality. To address these limitations, we propose **KVScout**, an efficient KV retrieval method that learns to select the KV entries needed for generating the upcoming sentence. With a lightweight retriever (8.9M parameters), KVScout builds a contextualized sentence index during prefill and queries it with the generation state before generating each sentence, enabling each retrieved KV window to be reused throughout that sentence. Extensive experiments on three long-context benchmarks demonstrate that KVScout maintains generation quality with a low refresh frequency, achieving 5.3× and 1.7–3.7× higher throughput than full-cache decoding and existing KV retrieval methods at 256K context, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.