Mask-Routed Sparse Attention for Long-Context Diffusion Language Models
Abstract
Diffusion language models apply bidirectional attention over the entire context at every denoising step, making each step substantially more expensive as context length grows. Sparsifying only at inference leaves the pretrained weights unchanged. We propose Mask-Routed Sparse Attention (MRSA), which further trains a diffusion language model to route using its native mask token. We append one mask to the end of each chunk and add a learned role offset to mark it as a routing slot. Every fourth layer uses a query with a bidirectional local window and the top-k chunks scored by that mask; all other layers use bidirectional sliding-window attention. Starting from Dream-v0-Base-7B extrapolated to 16K with YaRN, we compare MRSA with bidirectional NSA, DSA, and SWA under the same layout, key budget, and dense initialization, using the same Fast-dLLM decoder. Across three RULER probes from 16K to 128K, MRSA outperforms the other sparse methods when the YaRN factor is scaled to the evaluation length. At 128K, its mean accuracy is 18.4% higher than NSA, the next-best sparse model, while DSA and SWA remain near the floor. At the same length, decoding is 172× faster than full dense attention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.