Fill-and-Refine Decoding for Diffusion Large Language Models
Abstract
Diffusion large language models (dLLMs) have emerged as a promising alternative to autoregressive models. However, the vanilla decoding strategy suffers from a critical limitation: once an incorrect token is accepted, it cannot be revised, causing error propagation. Existing approaches perform refinement within the fill-up process, which limits their ability to exploit posterior information. In light of this, we propose FIRE, a two-stage decoding strategy consisting of a sequence fill-up stage followed by iterative refinement, where tokens are remasked and re-decoded conditioned on the remaining context. This mechanism allows previously accepted tokens to be reconsidered and corrected. Experiments on five benchmarks spanning language understanding, code generation, and mathematics demonstrate consistent improvements under identical computational budgets. These results highlight the importance of decoding algorithms in fully unlocking the potential of diffusion language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.