What Gets Committed First Shapes the Answer in Masked Diffusion Language Models
Abstract
Masked diffusion language models such as LLaDA and Dream generate text in parallel and in any order, and are now applied to retrieval-augmented question answering. At each step the model predicts many masked positions and the decoder selects which prediction to commit; the commitment is irreversible and becomes context for all later predictions. We ask whether the decoder's choice of which prediction to commit can redirect the final answer while the underlying model state stays fixed. Final accuracy cannot resolve this, since decodings that differ in one commitment differ in every later state. We therefore branch decodings from identical states and vary only the first non-padding commitment. This single change alters 27.3% of final answers. A dominant failure pattern arises when the decoder commits an end-of-turn (EOT) token before any answer content, while answer-supporting predictions are already present elsewhere in the canvas. Token identity governs the effect: committing end-of-sequence (EOS) yields 38.7% exact match at either tested position, and committing EOT yields 0%. Following this finding, EndFirst reweights EOS against non-EOS candidates once, at the earliest step where any uniform EOS reweighting can alter the trajectory, and returns control to the original decoder. With no extra forward pass, EndFirst at 32 passes outperforms a 64-pass retrieval-guided decoder on three QA benchmarks and improves that decoder by 5.3 exact-match points on average across eight LLaDA model-task settings. These results identify commitment selection as a distinct and directly controllable source of error in masked diffusion decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.