When Does Decoding Order Affect Coverage in Diffusion Language Models?
Abstract
Unlike an autoregressive language model, a diffusion language model can choose which positions to fill next. Prior work reports that forcing diffusion language models to decode from left to right can improve repeated-sampling coverage. For dLLMs, coverage (Pass@) matters because it shows how effectively repeated sampling finds successful solutions for test-time scaling and reinforcement-learning exploration. We ask whether left-to-right position order causes the gain. Our central insight is that changing order also changes which tokens commit together, what context they can use, when the decoder resolves uncertain choices, and how sampling temperature affects token choices. We hold model weights and the decoding loop fixed across three open diffusion language models. The left-to-right advantage is strongest with one token per step at low temperature, varies across model–benchmark pairs, and can disappear or reverse as width or temperature increases. At width two, left-to-right decoding commits neighboring tokens together, so the right-hand token cannot use its newly sampled predecessor. Avoiding these adjacent co-commits recovers much of the LLaDA/HumanEval coverage loss, with a compatible but less certain result on Dream. We also observe positive coverage changes without fixed left-to-right priority. Confidence-order decoding at a higher temperature reaches comparable coverage, while committing selected positions earlier yields weaker positive estimates. When we keep the number of early commitments roughly fixed, moving them later lowers coverage and raises single-sample accuracy. Thus, decoding order combines several decoder choices rather than changing position order alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.