From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion
Abstract
Uniform-state discrete diffusion updates all positions in parallel while keeping every token revisable, but earlier predictions are normally lost when the state changes. We study *persistent context*: *Committed* keeps selected clean-token predictions visible for up to the next model inputs while leaving the sampled state mutable. We ask two questions: why a denoiser should respond to the token shown at its own position, and how such a response affects the later sampling path. For the denoiser, we show that a class of cross-time distillation schemes, including Duo's consistency distillation, need not preserve leave-one-out (LOO) invariance even under an exact LOO teacher. On the sampling side, we show how a change caused by persistent context passes through the reverse updates and affects later states. This leads to an exact LOO example in which persistent context with a finite horizon improves the terminal distribution, while keeping the same context to the end makes it worse. On Duo and Duo-distilled, reference-token reinsertion produces opposite responses at moderate-to-high signal, and changing only the retained label shifts the loss–entropy tradeoff differently across checkpoints. Persistent input context also gives different loss–entropy and loss–repetition tradeoffs from using stored predictions in the reverse updates instead of the model inputs. Persistent context therefore does more than change the current sampling distribution: its effect depends both on how the denoiser uses the retained token and on how long that token remains visible during sampling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.