Disentangling Diffusion Decoding: Selection, Pace, and Numerical Semantics
Abstract
Masked diffusion decoders differ in position selection, commitment pace, verification, and numerical semantics, while scoring conventions also affect reported gains. We study these choices under explicit comparison contracts across five models, from three model families, and four benchmarks. In fixed-selector sweeps, doubling commitments per step halves denoiser evaluations and lowers accuracy in all 60 adjacent transitions. Selection gains also depend on pace. Across five controlled comparisons on GSM8K, confidence and two adapted attention-based selectors, ADAS and DAPD, are each regenerated at four and eight commitments under a shared harness and scorer within every checkpoint. On four models, ADAS's advantage over confidence increases 7.7–10.2 percentage points more than DAPD's as pace rises, with 95% confidence intervals excluding zero. These results characterize the adapted policies and do not establish task-general superiority. Replica batching changes 24–89% of raw output sequences across five models. We trace these output differences through termination, scorer normalization, answer extraction, and correctness. Extracted mathematical answers differ on 3.9–28.6% of items, whereas correctness changes on 2.4–7.0%; gains and losses partly offset, and every paired aggregate-accuracy interval contains zero. Supporting experiments show that token persistence does not prove peer reinforcement, while offline prediction and reference agreement do not guarantee effective or efficient decoding. Together, these results show why decoder comparisons must specify selection, pace, numerical execution, scoring, and verification rather than treating an end-to-end gain as a single intervention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.