acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Inefficient Training of Diffusion Language Models

Abstract

Diffusion language models (DLMs) need to be trained far longer than autoregressive models (ARMs) to reach the same held-out loss, despite their training objectives sharing a Bayes-optimal value. Proposed explanations include gradient noise, competition among corruptions and hard prediction problems; we test all three. We diagnose this training efficiency gap with a noisy quadratic model of training dynamics and find that it is largely explained not by gradient noise but by the features the DLM has learned, which permit far less loss reduction than an ARM's features taken at the same loss. This deficiency does not arise from averaging over corruptions, whose gradients point in the same direction, but from corruptions at high rates, which carry most of the DLM's loss in errors that its features can do little to correct. To distinguish limited expressivity from a failure to learn good features, we design Surface Sudoku, a Sudoku variant whose digits are latent variables. DLM training stalls at a loss far above an ARM's, while an auxiliary objective that supervises the latent digits helps close the efficiency gap. Our findings challenge the prevailing beliefs about the cause of DLM training inefficiency and suggest that closing the gap requires learning better features, through methods such as auxiliary objectives or adapting ARMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.