acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Mask Attention: Trainable Sparse Attention with Fused Block Screening

Abstract

Dynamic sparse attention must keep selection overhead low to translate sparsity into GPU speedups. We introduce Dynamic Mask Attention (DMA), an algorithm–systems co-design that fuses learned block screening into IO-aware attention kernels. DMA uses a rank-1 query-key gate to derive cheap blockwise bounds on gate scores, reducing memory traffic and bypassing QK score computation and post-score operations on skipped tiles. This screening is integrated into forward, backward, and decoding kernels without materializing a dense routing matrix. The same gate modulates retained attention logits and is trained through the language-modeling objective. Across models ranging from 80M to 8B parameters, DMA improves the quality-efficiency trade-off in pre-training, long-context evaluation, and sparse adaptation, while maintaining quality comparable to full attention. On Qwen3-8B, DMA achieves an average score of 64.26 versus 62.50 for Full Attention across RULER-128K, LongBench, and benchmarks of reasoning, coding, and knowledge. It also delivers approximately 1.2 the end-to-end prefill and decode serving throughput of Full Attention under an 8K-input/24K-output workload.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.