acceptodds
Under review as a conference paper at ICLR 2027

Learning When to Commit: Token Scores, Stopping Rules, and Stability in Diffusion Language Models

Abstract

Parallel diffusion decoding trades task accuracy for fewer backbone evaluations, and the accuracy retained depends on both token scoring and the rule that converts scores into commitments. We introduce a -parameter predictor of token stability, defined as agreement with a one-token-per-step decoding reference. Confidence determines candidate order, while a cumulative product of predicted stability scores controls how many tokens are committed, without additional backbone evaluations. Heads trained on at most problems are reused across tasks without retraining, with low measured inference overhead. Across three backbones and three benchmarks, the decoder has lower relative accuracy loss than its paired comparator in 71% of the measured comparisons at approximately matched mean cost. Two observations clarify what the learned score contributes. First, controlled comparisons show that a token score's observed accuracy advantage can depend on the stopping rule and decoding budget. Second, local stability and task accuracy differ. At approximately matched mean commitment counts, learned stability reduces reference disagreement relative to raw confidence and margin, but not significantly relative to a learned confidence-only control. Validation-selected evaluations and single-decision interventions do not establish consistent gains in final task accuracy. Our contribution combines lightweight, transferable stability-guided decoding with a controlled analysis of how token scores and stopping rules jointly shape the accuracy–cost trade-off.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.