Exploring the Boundary of Sparse Attention: Retrievability, Head Stability, and ANN Approximation
Abstract
Sparse attention is a promising way to reduce the cost of long-context inference, but it remains unclear when top- attention is accurate and whether practical approximate nearest neighbor (ANN) search can replace exact top- retrieval. We study these two questions separately—whether attention is retrievable by exact top-, and whether ANN indexes can find the keys it needs—across Llama-3.1-8B and Qwen3-8B on RULER, LongBench, and LongProc. Exact top- attention often approximates full attention well: across the 13 RULER tasks, the median number of keys needed to cover 90% of the attention mass grows far more slowly than the context, up to a model-specific length frontier. Its effectiveness nonetheless depends strongly on model, task, context length, and attention head. Retrieval-style tasks need significantly fewer keys than aggregation-style tasks, and fixed-budget top- attention degrades sharply at very long contexts, especially beyond a model's native context window. Head retrievability is structured and stable across inputs and decoding steps (held-out rank correlation ), and a head-aware budget calibrated once improves task scores by up to 7.3 points over a uniform budget of the same size. However, off-the-shelf ANN indexes such as IVF, HNSW, and PQ do not reliably recover enough attention mass under aggressive search budgets, and their set recall is not a reliable proxy for the mass they retain. These results suggest that sparse attention is useful but not universal: reliable deployment requires model- and task-aware strategies, adaptive budgets, and ANN methods tuned for attention mass rather than recall.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.