acceptodds
Under review as a conference paper at ICLR 2027

MSE Training Induces Diffusions That Are Nearly Flows

Abstract

In the DDPM paper, Ho, Jain, and Abbeel introduced two reversible diffusion processes parameterized by a noise schedule—a generator and an oracle process that the generator learns from—and derived a formula for the Kullback-Leibler divergence (KL) between these processes in the form of a time-weighted Mean Squared Error (MSE). However, they empirically found that omitting the weights improved performance on image synthesis benchmarks—a result later corroborated by many studies. More recently, removing the stochastic component during generation time has proven effective. (1) In this work, we provide a theoretical justification for these practices. We consider a broader class of diffusion processes (not necessarily reversible) parameterized by a noise schedule and a diffusion size b that share the same marginals. Since the weight associated with the MSE depends on b, omitting the weight is equivalent to solving the equation weight(b)=1, which yields a unique “MSE-diffusion”. For SOTA models, we checked that b is close to zero—the learned MSE-diffusion is nearly a flow—and we confirm this observation by comparing generators on ImageNet 512×512. These results suggest that flows beat reversible diffusions because training of SOTA models is an implementation of KL minimization for MSE-diffusions, which are nearly flows. (2) Moreover, we obtain a novel representation of the diffusion state as the sum of an explicit linear component, an unweighted pathwise integral of the denoiser, and a noise term. This representation offers the advantages of DPM-solvers while enabling the use of classical ODE methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.