acceptodds
Under review as a conference paper at ICLR 2027

Sample the Order, Not the Token: Test-Time Scaling for Masked Diffusion Language Models

Abstract

Standard test-time scaling methods like temperature-sampled self-consistency are highly effective for autoregressive models, but we show it is not true for masked diffusion language models (dLLMs). Attempting to scale dLLMs via temperature sampling collapses generation length and plummets accuracy (e.g., from 35% to 6% on MATH500), failing to yield the diversity required for ensembling. Motivated by this misalignment, we propose Order-Space Exploration (OSE), a novel decoding strategy demonstrating that dLLM diversity should originate from the order of token generation, not the token distribution itself. We first observe that a perfect model's answer distribution is invariant to unmasking order, thus shuffling this schedule yields "free" diversity driven entirely by the model's internal disagreement across generation paths, avoiding the accuracy penalty of temperature scaling. OSE introduces a one-parameter Plackett–Luce perturbation of the sampler's position ranking and aggregates the N draws with a majority vote over their answers. Empirically, OSE raises answer diversity by 2-5 while leaving single-draw accuracy unchanged. This unlocks test-time scaling for dLLMs, beating standard self-consistency on Dream by up to 45.5 points on a variety of downstream tasks. Additionally, OSE rollouts restore the learning signal in GRPO, removing nearly all zero-gradient steps on generation tasks and about half on multiple choice. It improves the trained policy where the base model has room to learn (+7.6 on MATH500, +8.5 on Countdown), with no change to the model or the inference cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.