Beyond Replacement: Levenshtein Editing for Diffusion Language Models
Abstract
Diffusion language models (DLMs) offer a flexible alternative to autoregressive models, enabling any-order infilling without specialized prompting. While most existing decoding methods treat decoded tokens as fixed, some recent models allow previously decoded tokens to be revised only by replacing them with alternative tokens, a paradigm we refer to as *replacement-only editing* (ROE). Although ROE can correct individual tokens, it remains inefficient for sequence shifts caused by missing or extra tokens and struggles when errors occur across many positions. To address these limitations, we propose LED (**L**evenshtein **E**diting for **D**iffusion Language Models), a post-training framework that extends ROE with insertion and deletion, enabling Levenshtein-style refinement of generated sequences. It uses two key training designs. First, it reconstructs edit supervision on the fly during supervised fine-tuning using longest-common-subsequence alignment, enabling the model to learn insertion and deletion operations without manually annotated labels. Second, it adopts multi-stage training, where the state from one stage becomes the input to the next, exposing the model to its own earlier errors and enabling it to learn how to correct them. When applied to the 16B-parameter LLaDA2.1-Mini Base, LED raises the average score across twelve benchmarks from to under confidence-guided decoding while increasing tokens per forward from to , whereas under confidence-free decoding it raises the average score from to . We further show that, in this setting, LED achieves higher pass@ at large than thresholded decoding without editing, demonstrating improved generation diversity and revealing the diversity potential of DLMs without confidence guidance. LED also improves robustness to large block sizes, raising the average score from to at block size . Together, these results suggest that LED supports more robust self-correction in DLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.