acceptodds
Under review as a conference paper at ICLR 2027

SubCenSA: Training-Free Block-Sparse Decoding with Count-Weighted Sub-Centroid Mass Estimation

Abstract

Block-sparse attention accelerates long-context decoding by reading only selected contiguous key-value (KV) blocks, while its quality and efficiency depend on identifying the blocks carrying most of the dense attention mass without scanning their original keys. We introduce Sub-Centroid Sparse Attention (SubCenSA), a training-free method that represents each key block with a few count-weighted sub-centroids, retaining relevant within-block structure at low storage and scoring cost. SubCenSA retains the full KV cache and block-aligned I/O, and computes standard softmax attention over the original keys and values in the selected read set. In exact arithmetic, each per-head estimate lower-bounds the true block mass and is at least as tight as mean-centroid pooling. For the shared grouped-query attention (GQA) selector, we show that the cross-atom oscillation of these residuals controls the self-normalized group-score error, yielding margin-based guarantees for recovery of the true group top-\(m\) blocks, retained group mass, and weighted fixed-state output error. We evaluate SubCenSA through selector-fidelity diagnostics, downstream performance on RULER, LongBench, and LongBench-V2-CoT, design ablations, and operator-level and end-to-end profiling at contexts up to 256K. SubCenSA improves over classical block indices by up to RULER points and achieves on LongBench and on LongBench-V2-CoT. At a 6.25% nominal complete-block KV budget, it remains at most of dense-attention quality while delivering up to operator speedups over dense FlashAttention-2.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.