SPEAR: Exact Self-Speculative Sampling from the Autoregressive-Order Policy of Diffusion Language Models
Abstract
Diffusion language models (dLLMs) can decode several tokens per forward pass and in any order, yet for reasoning the autoregressive (AR) order of the same model reaches a higher Pass@k. Sampling that AR-order policy exactly costs one forward pass per token, and the parallel decoders that avoid this cost sample a different distribution. We study whether exact sampling from the AR-order policy can be made cheaper on ordinary bidirectional dLLMs, and when it is worth its cost. We propose SPEAR, a training-free self-speculative sampler that drafts a block from one forward pass, verifies it against the exact AR-order conditionals in one batched pass, and keeps or resamples each token by the speculative sampling rule. We prove that its output follows the AR-order policy at any temperature. Experiments with three dLLMs on math and code show that the models agree with most of their own drafts, and that exact samples beat lossy parallel decoding at a matched number of forward passes. Matched on GPU time, the advantage disappears on math and survives on code only against a fixed-step decoder that breaks program syntax. Code is available at https://anonymous.4open.science/r/spear-code-iclr27-5C0B/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.