The Model Knows Where to Look: Distilling Attention Selectors for Long-Context Inference
Abstract
Language models now process whole documents, conversations and codebases, but generating each token requires reading a growing key-value cache. Attention often concentrates on a small fraction of that context, allowing sparse attention to reduce these reads. Selecting the relevant keys, however, introduces its own cost: the selector must inspect the cache or score a compressed representation of it. We present Highlight, a lightweight selector that skims the full context and highlights the relevant blocks for the model to attend to. It scores the context with a few narrow attention heads over its own small key cache, and selects 16-token blocks. It learns from the frozen model's own attention, through KL divergence on the tokens it selects. On five full-attention and hybrid models from 4B to 32B, accuracy stays close to dense attention on long-context benchmarks up to 128K tokens. On Qwen3-8B, every speedup we report comes from a configuration that stays within two points of dense or above it on each benchmark. At KV-cache saturation, selecting a quarter of the context yields 1.4× to 1.8× dense throughput. Adaptive top-p selection sets a budget for each query and layer, retaining 12 to 19 percent of the context and reaching up to 1.9× throughput. Quantizing the selector's key cache to four bits raises throughput up to 2.5× dense. Highlight also selects blocks up to 8 times faster than Twilight, an adaptive top-p selector, while both stay within two points of dense on held-out examples.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.