acceptodds
Under review as a conference paper at ICLR 2027

Unmasking Nondeterminism in Masked Diffusion Language Models

Abstract

Greedy decoding of masked diffusion language models (MDLMs) is nominally deterministic, yet the same prompt can yield different outputs when only the batch size or the GPU changes, while accuracy barely moves and so hides the problem. In autoregressive models, the first differing token permanently separates two generations, but MDLMs also decide at every step where to write, and how nondeterminism develops under this second decision is not understood. Tracking the partially decoded response over 512 steps for LLaDA and Dream, we find that forks arise mainly from the position decision. These forks are often transient: on LLaDA, 92.6% of prompts return to an identical intermediate response, and the token conflicts that make a difference permanent appear much later. Treating the two decisions separately, we freeze the position order under periodic LayerCast refresh and recheck only tokens within a few ULPs of a tie, cutting LLaDA's MATH500 output divergence from 69.2% to 1.0% at less than half the cost of LayerCast, and most of Dream's. Freezing changes the decoding rule, however, so the recipe's outputs agree with each other but often not with LayerCast's, and accuracy can shift. Reproducible MDLM evaluation should therefore track states beyond the first fork and report agreement across executions and accuracy separately. More broadly, locating where nondeterminism enters the decoding rule lets a remedy target the fragile decision instead of paying for full precision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.