acceptodds
Under review as a conference paper at ICLR 2027

Towards Solving the Unmasking Order in dLLMs via Constrained MaxEntropy Sampling

Abstract

A masked diffusion language model (dLLM) must choose which masked position to denoise at every step. Various rules exist for this choice, such as committing the most confident position (confidence order) and sampling positions in proportion to a tempered confidence (tempered selection). They are stated as separate procedures with no common objective, and tempered selection keeps every position eligible, so it cannot defer a commitment. We present **PACE**, a *position sampling strategy with adaptive confidence eligibility*, which keeps the odds of tempered selection but defers the least confident positions under an entropy budget relative to tempered selection. PACE follows from one constrained optimization problem: find the highest-entropy distribution over positions whose expected *damage*, the cost of a commitment, stays within a budget. Its solution is a Gibbs distribution over positions, and each rule we study corresponds to one choice of damage function, eligible position set and Lagrange multiplier, which is the inverse position temperature. A damage that depends on the committed token's probability and adds over successive commitments must be the surprisal, which gives the odds of tempered selection, and we prove that, among rules with these odds that defer only the least confident positions and meet the budget, PACE is the unique one that defers the fewest. We also show that selecting positions by the tempered confidence of their drawn tokens sharpens the distribution of written tokens toward an effective temperature that combines the token and position temperatures. On two model families and four benchmarks at token temperature , PACE improves single-sample accuracy on all eight model–benchmark pairs at the same compute, by to points over confidence order and to points over tempered selection, and leads at every number of samples up to ; its single-sample gain over tempered selection holds at a second token temperature.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.