From Denoising Dynamics to Endpoint Uncertainty in Diffusion Language Models
Abstract
Predictive uncertainty in stochastic language-model inference can be characterized by repeatedly sampling the same input and measuring the resulting distribution of outputs. However, doing so requires multiple complete generations. Masked diffusion language models expose an additional source of information: the evolution of the output state during denoising. We ask whether convergence observed within diffusion trajectories is systematically associated with uncertainty across their final outputs, which we refer to as endpoints. We study four masked diffusion language models on subsets of four reasoning benchmarks, collecting 20 stochastic generations per question for a total of 32,000 denoising trajectories. Within each trajectory, we measure within-run concentration, as the area under the curve describing the stable resolution of the eventual output over the denoising trajectory. Across repeated trajectories, we measure the uncertainty of the resulting endpoint distribution. Stronger within-run concentration is associated with lower endpoint entropy after accounting for model and dataset (, ). This association is robust to restricting the analysis to multi-token outputs and to reweighting the prevalence of positive-entropy cases. The relationship is strongly asymmetric: high within-run concentration almost exclusively accompanies concentrated endpoint distributions, whereas lower within-run concentration is observed under both concentrated and competing endpoint regimes. This relationship becomes increasingly apparent as denoising proceeds. Most importantly, features from a single denoising run predict whether independent held-out runs will concentrate on a common endpoint or produce competing endpoints. These results connect dynamics within individual denoising runs to variability across repeated inference. Within-trajectory dynamics contain information about whether independent repetitions of the same inference process will concentrate on a common endpoint or produce competing endpoints. This provides a concrete route towards estimating properties of the endpoint distribution with fewer complete generations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.