The Parallel-Decoding Score of Masked Diffusion Language Models Is Not a Mutual Information
Abstract
Training-free parallel decoders for masked diffusion language models decide how many positions to commit by bounding the dependence among them. The score they bound, , is read as a conditional mutual information, a reading that presumes the model's any-order conditionals are those of a single joint. We give an exact model-internal identity, : is the mutual information of the model's own one-at-a-time rollout and is a self-marginalization defect, the divergence between the marginal the model reports directly and the marginal its rollout induces. The identity splits any set's parallel error into a total correlation plus a sum of defects, and yields three one-sided certificates of incoherence. With estimators whose bias direction is known, the defect's mean is certified positive over pairs on four models and over scattered sets of two to eight positions on three, and on peaked anchors a conservative estimate puts it at a median 0.37–0.45 of the score. A set's parallel error also depends on the order in which it is revealed, by several times replicate noise. Where the two marginals disagree, the rollout's predicts the documents' own tokens better on average, so a parallel step samples from the less accurate one. Inside EB-Sampler's GSM8K decodes on LLaDA-8B-Instruct, half the decodes overrun the entropy allowance while the dependence it is written against stays within it: on the positions measured, the overrun comes from the defect. All code and the records behind every number are in the supplementary material and will be released at https://github.com/xxx/xxx upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.