Ordered-DLM: Ordering Token Generation through Syntactic Hierarchy
Abstract
Masked diffusion language models generate text by iteratively revealing masked tokens, enabling parallel generation and bidirectional conditioning. This flexibility raises a fundamental scheduling question: when multiple positions are masked, which token should be revealed first? Each reveal not only resolves one position but also changes the context available for predicting what remains. Motivated by information-gain principles in LLM reasoning, we prioritize revealing tokens that reduce uncertainty in predicting the remaining tokens, an objective we call *contextual planning*. To realize this objective, we use syntactic hierarchy to identify tokens with broad structural roles in an unsupervised manner. We introduce Ordered-DLM, a masked diffusion framework that predicts continuous syntactic heights, uses their relative ordering to induce differentiable hierarchy scores, and applies these scores to both forward masking and reverse denoising. The hierarchy is learned jointly with the diffusion model , requiring neither an external parser nor syntactic annotations. Experiments on LM1B and OpenWebText demonstrate lower perplexity than strong diffusion baselines on in-domain and zero-shot benchmarks, as well as lower generative perplexity across sampling budgets, while preserving parallel decoding. These results show that syntactic hierarchy-induced scheduling improves both contextual planning and language-modeling quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.