acceptodds
Under review as a conference paper at ICLR 2027

Soft Confidence Order in Masked Diffusion: Quality–Diversity Trade-offs Across Models

Abstract

Masked diffusion language models predict all masked tokens in one forward pass, leaving the decoder to decide which predictions to commit at each model call. Most-confident-first selection is common, but its strength is rarely isolated from other decoding choices at a fixed call budget. We study a one-parameter family of position-selection rules spanning confidence-free selection, soft confidence-weighted order, and hard most-confident-first ranking. We separate position selection from the time grid (tokens committed per call) and token revision. On two 124M checkpoints and LLaDA-8B-Base, we compare rules at equal call budgets using perplexity under a fixed GPT-2-large scorer, bigram diversity, and MAUVE. At temperature 1.0 and 32 calls on the 124M model, soft order is the best fixed-count rule tested: it lowers scorer perplexity by about 12% relative to both the standard ancestral sampler and a confidence-free fixed-count control, whereas hard ranking roughly doubles it. On LLaDA-8B-Base at the same temperature, the reduction relative to the ancestral sampler is 11–23% at 16–64 calls, with all 95% intervals excluding zero. On the 124M model, the gain over the ancestral sampler persists among outputs matched on per-document bigram diversity. One-shot revision trades diversity for a further 31% reduction in perplexity, while time grids fitted to pilot-run confidence lower ancestral-sampler perplexity by about 4% at the smallest budgets. Cooling to favors harder weighting. Decoding gains should be reported alongside call budget, count rule, and diversity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.