Reveal Policies in Diffusion Language Models
Abstract
Masked diffusion language models (DLMs) generate text by repeatedly predicting masked tokens and choosing which position to reveal next. We show that this reveal policy is part of the decoder specification: with the same conditional predictions, different causal policies induce different distributions over final answers. We call each such distribution a reachable answer law and their set the reachable answer-law polytope, characterize that polytope recursively, and determine what repeated samples under one or several policies identify. For a finite acyclic decoder with known transition kernels, Bellman and occupancy-measure programs exactly optimize terminal-answer utility under hard per-generation or expected model-evaluation budgets. We prove a sharp value-of-information result: response information beyond current per-position marginals can carry nearly half the utility range, establishing revealed history and lookahead as valuable signals for optimal reveal control. On trained models, with one token revealed after each of 256 evaluations, confidence ordering beats random ordering across four DLMs and matches or beats left-to-right decoding at the same budget. A four-path reveal-policy portfolio improves over confidence batching by 6.7 points for Dream-7B and 5.6 for LLaDA-1.5 on the GSM8K grade-school math benchmark. This demonstrates a matched-compute gain from combining reveal policies. On the Massive Multitask Language Understanding (MMLU) benchmark, 32 randomized-order continuations from a threshold-selected partial state end in its leading option in 95.6-99.2% of samples across four models. Thus threshold-selected partial states induce highly concentrated answer laws even under randomized future reveal order.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.