Flexible Reveal-Orders for Reinforcement Learning of Masked Diffusion Language Models
Abstract
Masked diffusion language models (MDLMs) have emerged as a promising alternative to autoregressive language models, motivating growing interest in reinforcement learning (RL) to improve their reasoning capabilities. RL for MDLMs, however, requires a tractable surrogate for the intractable order-agnostic likelihood, which marginalizes over a combinatorially large number of possible reveal orders. Existing approaches largely follow two strategies: ELBO-based methods optimize a tractable lower bound, while trajectory-based methods optimize the likelihood along a selected denoising trajectory. In this work, we provide a unified view of these approaches through the distribution over reveal orders underlying the log-likelihood gradient. Specifically, we show that the exact log-likelihood gradient is an average of per-order gradients under the posterior distribution over reveal orders, while ELBO-based and trajectory-based methods correspond to two extremes: a uniform distribution and a Dirac distribution concentrated on a single order, respectively. Motivated by this perspective, we introduce a flexible family of confidence-based distributions over reveal orders that interpolates between these extremes. Our approach can be integrated into existing RL methods without additional model evaluations. Across mathematical reasoning and planning tasks, we show that confidence-based order measures consistently improve both ELBO-based and trajectory-based RL baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.