acceptodds
Under review as a conference paper at ICLR 2027

Learning to Re-Draft: A Variational Stackelberg Game for Discrete Diffusion

Abstract

Discrete diffusion models offer the ability to re-draft, revisiting and correcting earlier tokens throughout generation. This capability depends on the forward corruption process that defines what the denoiser learns to correct. Masked diffusion models fix tokens once they are unmasked, while uniform diffusion permits revisions but relies on uniformly random token substitutions. We instead learn which token substitutions are most useful for training the denoiser to re-draft. We introduce Variational Stackelberg Discrete Diffusion (VSDD), a framework for learning for such a semantically aware corruption process. VSDD formulates training as a leader-follower game: the leader defines a Markovian corruption process parameterized by the denoiser’s token embeddings, while the follower optimizes a variational denoising objective with the noise process held fixed. The leader rewards corruptions based on how much the denoiser improves after learning from them, rather than on how easily they can be reconstructed by the current denoiser. We measure these improvements through a variational objective under a fixed reference corruption process. We approximate the follower’s response with a one-step gradient update and optimize the leader using a score-function estimator. We evaluate VSDD on molecular, text, and playlist generation domains. VSDD substantially improves molecular validity over uniform and masked diffusion, reduces text perplexity relative to uniform diffusion while remaining competitive with masked diffusion, and achieves sizeable improvements in offline playlist engagement metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.