Generation Order and Parallel Decoding in Masked Diffusion Models: An Information-Theoretic Perspective
Abstract
Masked diffusion models (MDMs) can accelerate sequence generation by decoding multiple positions in parallel, but two distributional questions remain central: how generation order interacts with model error, and what sampling bias is introduced by parallel factorization. We develop an information-theoretic framework separating order sensitivity from parallelization bias. Our analysis yields three insights. First, local conditional errors are evaluated under model-induced rollouts, so early mistakes can alter the contexts of later predictions. Second, factorized parallel decoding can produce individually likely but jointly implausible samples even with exact marginal predictions; forward KL, reverse KL, and incoherence reveal different aspects of this error. Third, exact correction by fixed whole-block rejection has expected proposal cost at least exponential in conditional total correlation. We study Prefix-Consistent Parallel Decoding (PCPD) as a practical relaxation that checks local prefix consistency rather than exact joint sampling. A controlled Block-HMM isolates sampling error, LLaDA provides ordering and parallelization diagnostics, and batched WeDLM experiments examine the quality–latency trade-off of verification. Together, the results explain why improving marginal predictions and correcting the sampling procedure address distinct obstacles to reliable parallel generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.