acceptodds
Under review as a conference paper at ICLR 2027

AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference

Abstract

Block sparse attention makes long-context decoding affordable by partitioning the KV cache into fixed-size blocks and loading only the blocks a query needs, and it has become the most deployable form of sparse attention because it leaves paged KV cache management untouched. Its weakness is accuracy: at a 4% KV budget, state-of-the-art methods trail full attention by up to 10 points on standard long context benchmarks. We trace a large part of this gap to a single design choice, the use of one block size for every attention head. Measuring attention recall head by head, we find that heads differ widely in their sensitivity to block granularity: some keep near-perfect recall, others lose over 90% of it when blocks grow from 16 to 32 tokens, and no single size serves both. We present AB-Sparse, a training-free algorithm–system co-design that gives each head its own block size. A lightweight one-time calibration assigns block sizes that transfer across domains; INT4 per-channel asymmetric quantization of block centroids, which only rank blocks and never enter the attention output, keeps centroid memory below that of the uniform baseline; and custom GPU kernels batch heads with heterogeneous block counts without padding and map variable-size blocks onto standard paged attention. On Llama-3.1-8B, Qwen3-8B, and Qwen3-32B, AB-Sparse raises the accuracy of Quest and ArkVale by 3.0–3.7 points on RULER and 1.9–2.6 points on LongBench on average (up to 5.4 points), recovering 27–85% of the gap to full attention, with comparable or higher decoding throughput.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.