acceptodds
Under review as a conference paper at ICLR 2027

Watch Your (Time-)Step: Adaptive Sampling Helps Train Your Diffusion Language Model

Abstract

The success of _auto-regressive language models_ has spurred interest in _diffusion language models_, which can generate multiple tokens in parallel. Recent masked-diffusion language models often reuse bidirectional denoisers familiar from masked language modeling, but differ in how they distribute training across corruption levels. We focus on one common assumption in these models: uniform time-step sampling for noising sequences. We first view _masked diffusion language modeling_ (MDLM) as risk estimation over a continuum of fixed-rate _masked language modeling_ (MLM) denoising objectives. This perspective explains both why a fixed masking rate can be optimization-friendly, since it repeatedly trains the same corruption level, and why MDLM requires ELBO-dependent weights to estimate the full diffusion objective without bias. The same analysis shows that weighted losses can have time-step-dependent second moments, making uniform time-step sampling statistically inefficient. Motivated by this, we propose _**No**rmalized **2**nd **Mo**ment **Aware**ness_ (**NoNoMoAware**) sampling, an adaptive strategy that maintains online estimates of the weighted-loss second moment, fits a proposal over time-steps, and corrects with importance weights to preserve unbiased loss estimation. For empirical validation, we conduct experiments using different backbones during both pre-training and instruction-tuning, with our method showing downstream improvements on a number of different benchmarks. Our findings illustrate the need for more principled investigation into the training dynamics of DLMs and offer practical guidance for training such models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.