Equilibrium Diffusion Language Models
Abstract
Diffusion language models generate text through iterative, parallel token updates. When applying these models to reasoning tasks, we often seek any correct solution, rather than reproduce the distribution of answers in the training data. We introduce *Equilibrium Diffusion Language Models*, which approach generation as a process of finding and preserving valid solutions. We train the model with a single cross-entropy objective to reconstruct valid solutions from corrupted inputs. Our theoretical analysis identifies conditions on the corruption process under which valid solutions become absorbing states of the population-optimal model and repeated sampling reaches them almost surely, with exponentially fast convergence. The analysis also establishes a probability gap between preserving valid and invalid states, suggesting a criterion for stopping generation. We evaluate the approach on unconditional generation of valid Sudoku grids and program synthesis with training on TinyGSM and evaluation on GSM8K. On GSM8K, our model achieves 63.6% accuracy, comparable to 63.3% for the autoregressive greedy baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.