acceptodds
Under review as a conference paper at ICLR 2027

SPECULATIVE CORRECTION: DRAFT-THEN-REFINE DECODING FOR DIFFUSION LANGUAGE MODELS

Abstract

Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text block by block. We show that existing pretrained DLMs can directly refine previously completed block-autoregressive responses, without additional training or architectural modification. This yields a simple draft-then-refine decoding pattern: a model first generates a complete response, then treats that response as an editable initialization for bidirectional refinement. We study both same-model correction, where one checkpoint drafts and refines, and small-to-large speculative correction, where separately released small and large checkpoints are composed without joint training or adaptation. Same-model correction improves the evaluated quality-latency trade-off: on a held-out GSM8K set, it raises accuracy from to with a -token budget while running faster, and on MATH-768 it matches the Flash score of while running faster. Small-to-large correction retains much of the cheaper model's speed. On MATH-384, it scores rather than while running faster than Flash. On MBPP+, it raises the fraction of programs passing the expanded tests from to while running faster in standardized timing. Refinement is sparse: it changes only about five to six final tokens per response and fixes substantially more answers than it harms. Taken together, these results show that a completed response provides enough structure for a short refinement stage to make useful changes. Within this plug-and-play framework, same-model correction shows that an existing checkpoint can achieve a better quality-latency trade-off through decoding alone, while small-to-large correction can combine separately released checkpoints into a cascade that often adds an intermediate operating point between them. The success of this off-the-shelf cascade further motivates training small and large DLMs for complementary drafting and refinement roles as a dedicated architecture.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.