acceptodds
Under review as a conference paper at ICLR 2027

Stratified Sparse Attention: Reducing the Sparse-to-Dense Gap at Extreme Sparsity

Abstract

Large Language Model inference over long contexts is increasingly bottlenecked by the cost of attending to a growing Key-Value (KV) cache. Sparse attention reduces this cost by accessing only a fraction of the cache, but existing methods face a trade-off between deterministic retrieval, which captures high-scoring tokens but discards residual attention mass, and sampling, which improves coverage at the cost of higher variance. Hybrid methods combine top- retrieval with uniform sampling, yet still exhibit a substantial gap from dense attention at aggressive sparsity. We introduce *Stratified Sparse Attention (SSA)*, a unified framework combining deterministic retrieval with score-guided sampling. SSA deterministically attends to high-scoring tokens while sampling the residual using proxy attention scores and correcting for their inclusion probabilities, thereby recovering attention mass beyond the retrieval set. SSA encompasses top-, top-, and hybrid top-+sampling as special cases. Across RULER, LongBench, LOFT, and AIME, and five models ranging from 3B to 27B parameters, SSA consistently improves over the baselines at matched sparsity. On RULER-HARD, SSA improves accuracy over the strongest baseline from 72.5% to 83.7% at 32K and from 69% to 76.8% at 128K over the strongest baseline under sparsity, recovering 94% and 91% of dense-attention performance, respectively. With optimized CUDA kernels, SSA achieves up to decode speedup over FlashInfer at 1M context, demonstrating that its accuracy gains translate into practical long-context inference acceleration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.