acceptodds
Under review as a conference paper at ICLR 2027

Better Selector, Better Decoder? Evaluating Token Selection in Diffusion Language Models

Abstract

Parallel diffusion decoding can accelerate text generation by filling several masked positions per forward pass. Its accuracy depends on both which positions are selected and how many are filled at once. Adaptive decoders couple these decisions, because earlier selections shape the context used to determine later counts. Changing a selector can therefore also change its schedule of per-step counts, while fixing the schedule leaves unclear whether the measured advantage generalizes. We record the per-step counts that adaptive decoders choose and let different selectors fill exactly these counts. As a control, we spread the same positions evenly over the same number of forward passes. On GSM8K in both LLaDA and Dream, DAPD’s selector is more accurate than confidence when both use DAPD’s counts. By contrast, confidence is more accurate when both use counts from a confidence-threshold decoder. We also let a different selector fill the first half of the output positions. The original decoder then completes the answer more accurately by choosing new counts than by reusing its recorded counts. These findings motivate the Decision-Value Protocol. It recommends evaluating a selector together with the rule that sets its counts, choosing counts from the current text, and choosing complete decoders by accuracy and measured latency. In a later study on ASDiv, following the protocol selects a decoder that is 9.38 points more accurate than the one a selector comparison favors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.