ReFresh: Learning When to Retrieve for Long-Context Prefilling
Abstract
Long-context language models enable retrieval and reasoning over large documents, but quadratic self-attention makes prefilling a major bottleneck. Hybrid attention mitigates this cost by combining a few full-attention heads with cheaper streaming heads, yet its efficiency remains limited by the full-attention budget, whose reduction can hurt long-range retrieval. We observe that full attention serves two roles—computing attention outputs and identifying relevant context regions—and that the latter exhibits substantial redundancy across layers: the same key/value (KV) head in neighboring layers often attends to highly overlapping context regions. Existing cross-layer reuse methods exploit this redundancy, but typically reuse indices or attention according to fixed or heuristic schedules. Building on this observation, we introduce ReFresh, a hybrid attention framework that learns where global retrieval should be refreshed or reused across layers, yielding a binary Anchor/Reuse map through dense-to-sparse distillation with the language model frozen. Anchor slots perform full attention, derive block scores from the resulting attention probabilities, and cache the selected KV-block indices for each query block. Reuse slots attend sparsely using the latest indices from an earlier Anchor of the same KV head, preserving global access to distant blocks without recomputing selection—unlike streaming heads, whose access is restricted to attention sinks and a recent window. Across Qwen3-8B and Llama-3.1-8B on LongBench, LongBench v2, and RULER, ReFresh preserves long-context quality with 80% Reuse and achieves a prefill speedup for Qwen3-8B at 256K tokens over FlashAttention-2.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.