Commit Order, Not Commit Rate: Answer Availability and Prefix Preservation in Diffusion Language Models
Abstract
Diffusion language models commit multiple tokens in parallel, but their decoding configurations are often treated as implementation details when safety and utility are evaluated. We show that this assumption is unsafe: commit order (which positions may resolve together) can matter more to answer delivery than commit rate (the number of tokens committed per forward pass). We audit identical weights under paired changes to eligibility window and confidence-threshold commitment, separating substantive answers, semantic outcomes, and inference cost. Widening a blockwise window degrades answer availability on boundary requests even when threshold commitment is retained; a full-canvas, one-token arm isolates the window effect but is not a deployed decoder. A per-request state ledger shows why a lower all-request harmful-compliance rate can reflect lost answers rather than safer content. We then preserve the first block of a reference decoder exactly before parallel suffix commitment. Same-forward-budget position controls are consistent with a position-specific effect, and paired adjudication finds higher benign-boundary helpfulness than blockwise threshold decoding at higher latency. These results identify a quality-oriented operating point while harmful-compliance differences remain unresolved.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.