DUEL: Exact Likelihood for Masked Diffusion Reveals Improved Pre-Training Scaling Laws
Abstract
Masked diffusion models (MDMs) are trained to predict tokens in every order but generate in one. As such, MDM likelihood is an intractable sum over all orders. The evidence lower bound (ELBO) is used as a proxy, but it is a loose bound and assumes positions are unmasked uniformly at random. Frontier MDMs, however, select positions with deterministic rules such as confidence. To resolve this mismatch, we introduce DUEL, a framework that unifies common deterministic MDM samplers. DUEL allows for likelihood computation that is both exact and under the sampler that is actually used: replay generation, reveal the true tokens instead of sampled ones, and sum their log-probabilities. DUEL is a simple plug-and-play algorithm applicable to any pretrained MDM with no retraining. Evaluating with DUEL instead of the ELBO changes what we conclude about MDMs. The perplexity gap between MDMs and autoregressive (AR) models shrinks by up to 32% in-domain and 82% zero-shot. In scaling laws, the reported MDM–AR compute gap is under DUEL: between and of it was an evaluation artifact. Scoring multiple-choice answers with DUEL improves LLaDA-8B accuracy on all six tasks, by up to 9.7 points. DUEL also selects samplers correctly: it matches a gold-standard LLM judge on every decided comparison, which generative perplexity and MAUVE fail to do. DUEL reveals that MDMs are ultimately more competitive with AR models than the ELBO suggests.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.