acceptodds
Under review as a conference paper at ICLR 2027

SPIRA: Sparse Pre-RoPE Informed Routing Attention for Diffusion Language Models

Abstract

Block diffusion language models generate a block of tokens at a time by denoising masked positions in parallel. Sparse prefix attention reduces the number of prefix keys attended to during decoding, but selecting those keys is itself a full-prefix operation when each masked query scores the prefix independently. We find that the masked queries of a block share a common direction before RoPE. This structure supports shared candidate generation across the block, while final selection still depends on query-specific scores from the original position-dependent queries. SPIRA exploits this separation: one routing query per query head, constructed from the block’s pre-RoPE query mean, scans the prefix once per block to retrieve a candidate set, and the original queries rescore only those candidates to select the final keys. The selected keys are then reused throughout the block, so later denoising steps no longer read the full prefix. Across three block diffusion language models, SPIRA outperforms two sparse baselines on LongBench with 256 retained keys, exceeding the stronger baseline by up to 10.2 points. With 1024 keys, it comes within 0.06 points of exact attention on SDAR-4B-Chat-b64. On the same model, it achieves a 3.45× speedup in per-block decode latency at a 122k-token prefix.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.