Dependency-Structured Generation with Masked Diffusion Language Models
Abstract
Autoregressive language models generate along a fixed left-to-right chain, whereas standard masked diffusion language models (MDLMs) train on randomly masked contexts, assigning equal training weight to all possible token reveal orders. Between these endpoints lies a structured regime: several tokens may be ready concurrently, while others should wait for specific evidence. We formalize this regime as dependency-structured generation, in which a model’s reveal behavior is compactly described by a conceptual directed acyclic graph (DAG) over token positions, whose edges encode prerequisite relations. The framework recovers left-to-right autoregressive and standard (blockwise) MDLM objectives as boundary cases. The graph is a behavioral abstraction rather than a generated variable and is neither supplied nor constructed at inference. We instantiate this framework with MAPLE. During training, MAPLE uses a frozen teacher (e.g., a pretrained LLM) to extract a dependency graph for each target sequence; this graph induces a normalized training distribution over reveal orders. We derive an exact negative ELBO that jointly models the next reveal position and its token value. Under standard realizability, every global minimizer of the population negative-ELBO objective recovers the data sequence distribution. This optimality guarantee does not require the learned model to exactly reproduce the teacher’s reveal behavior; the extracted dependencies instead serve as a finite-model inductive bias toward informative partial contexts. At inference, MAPLE retains the original parallel MDLM decoder unchanged. Continued pretraining of LLaDA 2 and DreamReasoner for mathematical reasoning and code generation improves task accuracy and substantially reduces teacher-dependency violations without sacrificing decoding parallelism.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.