Closing Performance Gap for Diffusion Language Models under Low Inference Budgets
Abstract
Diffusion language models offer an architectural route to low-budget generation by updating multiple tokens in parallel, which can reduce the number of sequential decoding steps relative to token-by-token autoregressive decoding. However, realizing this advantage at low budgets is challenging, as generation quality often declines when the number of denoising and refinement steps is reduced. Editable diffusion models address this quality loss through post-draft refinement, but useful intermediate results may still be lost before final submission. We introduce CaRR (Commitment-aware Refinement and Recovery), a training-free framework that couples token remasking with trajectory-based output recovery. Token Commitment Remasking (TCR) remasks selected risky pure replacements and refills them on the next scheduled edit forward using the updated context. Submission Commitment Recovery (SCR) uses the resulting refinement trajectory to recover supported candidates when the final response does not satisfy task-specific output requirements. Both modules operate within the native mask/edit schedule without additional backbone forwards. Experiments on code generation and mathematical reasoning show consistent gains in average accuracy over native ME-DLM across four inference budgets, with the largest gains under tight budgets. In particular, at one-eighth of the full scheduled forward budget, CaRR improves the average accuracy across six metrics over native ME-DLM by 7.44 percentage points, including gains of 11.37 and 10.05 points on MBPP and MBPP+.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.