acceptodds
Under review as a conference paper at ICLR 2027

Scaling Order Learning for Diffusion LLMs

Abstract

The standard masked diffusion model (MDM) objective treats all unmasking orders uniformly, potentially allocating substantial training compute to orders poorly suited to individual examples. The resulting training inefficiency makes high-performing MDMs costly to train even when initialized from strong autoregressive models (ARMs). We introduce SolDiff, a scalable variational framework that learns data-dependent order distributions to train stronger MDMs under limited compute budgets. We model orders as latent permutations, with a learned posterior focusing token-prediction training on mask patterns induced by preferred orders and a learned prior reproducing these preferences at inference. To scale this formulation to MDM pretraining, we systematically identify key design choices for stable and effective optimization and introduce rank-weighted auxiliary objectives to support parallel decoding. Across 12 benchmarks, SolDiff-8B outperforms all other evaluated ARMs and MDMs of comparable size on average. We highlight that this result requires only 100B tokens of continued pretraining, less than 0.3% of the base model's pretraining tokens and substantially fewer than those used in existing ARM-to-MDM adaptations. In controlled comparisons, SolDiff outperforms standard block diffusion trained with as much compute. We further observe consistent gains across model sizes from 1.7B to 14B.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.