acceptodds
Under review as a conference paper at ICLR 2027

Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models

Abstract

While reasoning on autoregressive (AR) models is often performed by chain-of- thought reasoning and reflection, their refinement of previous outputs still relies on fully sequential generation, even when only local edits are needed. In contrast, the masking mechanism in Mask Diffusion Models (MDMs) naturally supports explicit local edits on previous outputs, allowing selective refinement without discarding previous answers and generating another from scratch. While this property more closely aligns with how humans correct mistakes by iterative local refinement, standard MDM decoding does not allow previously revealed tokens to be actively revisited and re-masked for further revision. We propose Reflective Masking (RM), a lightweight post-training framework that equips MDMs with context-dependent iterative self-correction. RM enables a model to revisit and revise prior outputs based on the evolving context, providing an adaptive way to allocate additional inference computation to selective revision rather than forward generation alone. To exploit information from previous turns, we further introduce History Reference, a parameter-free mechanism that incorporates intermediate denoising states during revision. Our approach requires no architectural changes and is easily applicable to existing MDMs. Across diverse tasks including text generation, Sudoku, and image editing, Reflective Masking consistently outperforms standard masking-based baselines, suggesting that explicit in-place revision can serve as a useful primitive for iterative reasoning and self-correction in MDMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.