acceptodds
Under review as a conference paper at ICLR 2027

When Does State-Dependent Denoising Preserve Its Target?

Abstract

Diffusion language models adaptively choose which masked tokens to reveal, revise, or remask from the current model state. Even when every local update uses a valid conditional distribution, their state-dependent composition can change the distribution of completed sequences. We formalize this gap as a target-transport problem and identify two distinct boundaries. Specifically, for monotone generation, any predictable sequence of exact singleton or joint-block conditionals preserves the target, whereas independent parallel block updates are exact only when the selected conditional factorizes. For revisable generation, a state-dependent mixture of target-reversible kernels preserves the target exactly when its selector commutator vanishes. An action-ratio acceptance rule restores invariance for arbitrary positive selectors under reversible component kernels, while leaving within-block factorization error untouched. Empirically, exact finite systems verify these identities and exhibit total-variation defects up to 0.79 outside the target-preserving regimes. Normalized neural language targets reproduce the same separation: experiments with frozen LLaDA and Dream models further show endpoint divergence that grows with revision horizon, selector concentration, and remasking width, while state-independent controls remain matched. Together, these results distinguish routing error from block approximation and provide explicit criteria for target-preserving adaptive denoising.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.