Beyond Autoregressive Recipes: Scaling Up Mixture-of-Experts Diffusion LLMs to 30B
Abstract
Diffusion large language models (dLLMs) are a promising alternative to autoregressive (AR) models, but existing sparse Mixture-of-Experts (MoE) dLLMs largely inherit AR recipes, and the masked denoising objective calls for dLLM-specific scaling laws. We present a systematic study of scaling laws for MoE dLLMs, spanning optimization hyperparameters, compute allocation, and MoE architecture across a broad range of compute budgets. Guided by these laws, we train LLaDA MoE v2, a 30B-A3B dLLM, from scratch on 23.5T tokens; it surpasses the previous work LLaDA MoE at only half of the latter's training FLOPs. The base model approaches the strong open-weight AR MoE model Qwen3 30B-A3B on knowledge, reasoning, and coding benchmarks while using \(65%\) as many pretraining tokens. After supervised fine-tuning, it outperforms SDAR Chat, a competitive MoE dLLM, on seven of eight reasoning and coding benchmarks. Together, these results show LLaDA MoE v2 to be the strongest dLLM trained from scratch to date, establishing dLLM-specific scaling laws as a foundation for large-scale MoE dLLM training and practical application.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.