ConGO: Consistency-guided Optimization for Parallel Decoding in Discrete Diffusion
Abstract
Discrete diffusion is an emerging paradigm for generative AI, in which a sequence of tokens is generated through iterative updates using a learned denoising model. While the standard formulation randomly selects token positions to update in each denoising iteration, common heuristics select the positions with highest marginal confidence or minimum inter-token dependency to increase generation quality. These approaches, however, strictly isolate the position selection from the content selection. Such *value-blind* position selection hinders the commitment of jointly consistent tokens, which can be indicated by explicit constraints or learned priors. In this work, we propose consistency-guided optimization (ConGO) for parallel decoding in discrete diffusion models, which formulates the update selection as a quadratic unconstrained binary optimization problem. Guided by a lightweight, *value-aware consistency model*, ConGO offers a flexible spectrum of integration depending on the reliability of the available knowledge. When strong, explicit domain knowledge is available, such as in Sudoku, ConGO *jointly* selects both the positions and their content, achieving significantly higher accuracy on hard problems (up to a 42.4 pp improvement). When explicit constraints are absent, ConGO restricts its intervention purely to position selection based on learned pairwise priors (e.g., bigrams). Integrated into both marginal and dependency-aware decoding of LLaDA-8B-Instruct in coding, ConGO improves each method's accuracy-NFE trade-off towards fewer function evaluations, with the dependency-aware integration establishing a novel Pareto front.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.