acceptodds
Under review as a conference paper at ICLR 2027

Learning from Failures: Verified Edit Distillation for Variable-Length Diffusion Language Models

Abstract

Variable-length diffusion language models insert and delete tokens as they generate, so their outputs are not tied to a fixed set of positions. This matters when the model is to learn from its own failed output. A corrected version of a failure usually differs from it in length, so the two do not share positions, and the standard masked SFT objective, which supervises tokens at fixed positions, can only treat the correction as a whole new response; it cannot express which tokens of the failure to keep and which to change. We propose Verified Edit Distillation (VEDI), which obtains such corrections as execution-verified teacher repairs of the model's own failures. VEDI aligns each failure with its repair, accounting for insertions and deletions, and jointly trains one backbone to predict where to edit and to reconstruct only the edited region given the preserved context. Inference uses the original sampler without a teacher or an additional repair stage. On two 7B models, FlexMDM and DreamOn, VEDI improves over Base on all five benchmarks, with gains of up to 12.7 HumanEval+ points on DreamOn. Trained on the same failures and repairs, full-rewrite SFT recovers only part of this gain: VEDI exceeds it by 6.3 and 5.0 HumanEval+ points on FlexMDM and DreamOn and matches or exceeds it on MBPP+. Controlled comparisons show that the gain requires a real student failure, a repair written for that failure rather than a reference solution, and edit-location supervision that reaches the shared backbone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.