Decentralized Diffusion Language Models
Abstract
Discrete diffusion language models (DLMs) generate text by iteratively denoising masked sequences, and their training objective generalizes autoregressive modeling beyond a predetermined generation order. Like other large language models, however, DLMs are trained as monolithic models, whose data-parallel scaling across GPU nodes is limited by cross-node synchronization and by diminishing returns from larger batches. We propose Decentralized Diffusion Language Models (D-DLM), which partition the training corpus into disjoint clusters, train one expert DLM per cluster without cross-expert synchronization, and train a router to predict the cluster membership of a noisy sequence. We prove that the probability velocity of the DLM reverse process decomposes into a mixture of cluster-conditional velocities weighted by the cluster posterior, and that the clean-token posterior admits the same decomposition. Hence, in the realizable limit, D-DLM recovers the same reverse process as centralized DLM sampling, and under masked diffusion, sampling reduces to a router-weighted ensemble of expert predictions. In pre-training at up to 1.2B parameters and 80B tokens, D-DLM improves the average downstream accuracy over monolithic training at every scale, by up to 2.05 points, with no cross-node communication, and it improves training throughput by up to 24% when training is bottlenecked by cross-node communication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.