acceptodds
Under review as a conference paper at ICLR 2027

When ODE Fidelity Misleads Diffusion Timestep Search

Abstract

Diffusion models can generate images by solving an ordinary differential equation (ODE) that transforms noise into an image. A numerical solver follows this process through repeated model evaluations, using each prediction to update the current state. Fast samplers use fewer evaluations to reduce computation. A schedule specifies when these evaluations occur. We examine choosing among candidate schedules by minimizing the distance from the final state to a reference computed with many more steps. We change the evaluation times while keeping the model, sampling procedure, and total number of evaluations the same. The selected schedules can bring the final state closer to the reference yet score worse on automated preference metrics and on metrics comparing generated and real image distributions. Under strong classifier-free guidance, which increases the text description's influence, the search can leave long gaps in the low-noise region near the end of generation. For a schedule with such gaps, smaller steps early in generation reduce the distance to the reference more, whereas smaller steps near the end improve the learned image-preference score PickScore more. Using fewer evaluations early and more near the end can improve PickScore without increasing the total computation, although the distance to the reference increases. Constrained endpoint search limits the largest low-noise gap and recovers most of the PickScore loss in the tested candidate sets on Stable Diffusion 1.5, Stable Diffusion XL (SDXL), and PixArt-Σ. The failure lies in how the search allocates computation: evaluation times that improve numerical accuracy can leave too few evaluations where they matter more for image quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.