acceptodds
Under review as a conference paper at ICLR 2027

When to Revise: A Scaling Study of Masked Diffusion Language Models

Abstract

Training-free remasking overlays let masked diffusion language models (MDMs) return committed tokens to the mask and decode them again. Each overlay is evaluated in its own harness and usually at one operating point, so it is unclear whether a reported gain reflects the method or the step budget, the compute spent, and the sample of benchmark problems. We study this question for four overlays (CoRe, ReMDM-conf, STaRR and T2M) and vanilla decoding on LLaDA-8B-Base with a common-harness protocol. It runs all methods in one decoder at three step budgets, counts each method's forward passes from its implementation, adds equal-cost vanilla controls, and estimates paired uncertainty for the equal-weight benchmark mean. CoRe, the only overlay that spends extra forward passes to verify tokens, gains about two to three points over vanilla at every budget, and the gain does not shrink under equal-cost controls. Its lead over the best of the three compute-free overlays, however, is within benchmark sampling uncertainty at two of three budgets. The overlays differ far more in how many answers they change than in net accuracy, and on a second backbone CoRe's effect on one benchmark reverses sign. Code is available at https://anonymous.4open.science/r/dLMM-benchmark-iclr27-81D2/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.