acceptodds
Under review as a conference paper at ICLR 2027

Correction Training for Diffusion Models

Abstract

Neural networks inevitably produce imperfect predictions, and accumulate errors in iterative sampling scenarios like generative diffusion models. In this paper, we introduce an iterative correction step for diffusion models. The corrector takes the base model’s own prediction as input and is trained against the same ground-truth target as the base model. This makes correction an explicit learning task, with training inputs provided by the base model. During sampling, at each denoising step, we iteratively refine the denoiser’s clean-sample prediction. Although the corrector is trained only on predictions from the base denoiser, we find that it can be applied repeatedly to its own outputs at inference time for even better results. This simple framework substantially improves sample quality while largely preserving diversity. We extensively evaluate our method across widely used architectures and training frameworks, including DiT, SiT, REPA, and JiT. Our method substantially improves upon previously reported FID scores for unguided generation and, notably, unconditional generation, where producing high-quality samples without guidance has been particularly challenging. Our method achieves performance competitive with classifier-free guidance (CFG), and combining correction with CFG further improves diversity and reduces oversaturation typical of CFG. Finally, we show that the additional inference cost of correction can be amortized by distilling the refined predictions back into the base diffusion model, retaining the quality gains without additional inference-time overhead. We hope that our findings encourage further efforts to move beyond diffusion models' long-standing reliance on guidance techniques.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.