acceptodds
Under review as a conference paper at ICLR 2027

ReadSpan: Anticipating Where Decoding Reads for KV Cache Eviction

Abstract

Key-value (KV) cache eviction reduces the memory of long-context inference by keeping only the input positions that an observation window at the end of the input attends to. This snapshot misses a recurring pattern of grounded responses: decoding jumps from the cue that the window attends to onto a nearby entry point and then reads forward through a span that the window barely sees and eviction removes first. We propose \method, a training-free method that anticipates this span from the first output token, which prefill predicts before any eviction. On top of a head-matched base selection, \method routes the window scores to the positions holding the first output token, extends them forward, and shifts each head's budget toward the span only to the extent that the base selection misses it, without any extra forward pass. On LongBench and RULER with Llama-3.1-8B-Instruct and Qwen3-4B, \method outperforms six training-free baselines on the combined score in all four model–budget settings; at 5% KV cache, it raises the RULER average over the strongest baseline by 3.4 and 7.2 points.\newline Code and data: https://anonymous.4open.science/r/ReadSpan-5443anonymous.4open.science/r/ReadSpan-5443.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.