ReScope: Sparse Prefill with Segment Observations and Query Updates
Abstract
Long-context language models incur substantial prefill costs as dense attention scales quadratically with input length. Sparse attention reduces this cost by selecting query–key interactions, but a few representative queries can miss important interactions when tasks and evidence are distributed throughout the input. Expanding observations, meanwhile, increases selection overhead. We propose ReScope, a training-free sparse prefill method that combines segment observations with temporary query updates. Segment observations cover different input regions, while a small tail window repeatedly reads the context and updates temporary query states to identify additional candidate interactions. Keys, values, and positions remain fixed during these updates. Final sparse attention uses the original queries, keys, and values and retains the full KV cache. On evaluation workloads constructed from real question-answering data and natural documents, including HotpotQA, NarrativeQA, and 128K scientific-paper question answering, ReScope improves F1 by 3.0–4.9 points over the respective sparse baselines. On a controlled retrieval benchmark with 128K-token inputs, optimized ReScope achieves 100% accuracy and a 2.45× complete-prefill speedup over dense attention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.