SABER: State-Aware Budget Estimation and Routing for Sparse Attention
Abstract
Sparse attention accelerates long-context Transformer inference by selecting a subset of key–value (KV) entries for each query. Existing adaptive sparse-attention methods typically rely on posterior pruning, leaving the upstream pre-selection budget as a manual design choice and introducing irregular computation patterns. We propose State-Aware Budget Estimation and Routing (SABER), an internal-probe-guided framework that unifies pre-allocation and online control of sparse-attention budgets. SABER predicts, for each query or generation state, the sufficiency of candidate budgets using lightweight probes over intermediate hidden states of the language model, and allocates the smallest budget that preserves task performance. This mechanism applies consistently across generation: it selects sufficient budgets from prompt states and, for long reasoning, continuously refines them as the generation trajectory evolves, enabling both one-shot allocation and trajectory-aware dynamic adjustment in a single framework. By performing input- and state-dependent allocation before and during decoding, SABER avoids over-provisioned attention and reduces unnecessary computation. Across long-context and reasoning benchmarks, SABER preserves accuracy relative to strong fixed-budget, top-, and trainable sparse-attention baselines while reducing average selector budget by up to 67.6% and achieving up to 1.35 end-to-end speedup over state-of-the-art fixed-budget sparse attention methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.