acceptodds
Under review as a conference paper at ICLR 2027

ENCORE: Training-Free Few-Step Diffusion Sampling with Terminal Moment Matching

Abstract

Diffusion models achieve state-of-the-art quality in visual and audio generation, but sampling requires many evaluations of an expensive denoising network, and fewer steps severely degrade sample quality. Existing approaches either design fast ODE solvers or rely on auxiliary trained networks for post-hoc refinement. We show that even with optimized ODE solvers, few-step generation deviates systematically from many-step runs in two distributional ways: (i) a loss of power across all frequency bands and (ii) a mean shift in a feature space. Because both are distribution-level properties, they cannot be precisely measured or corrected along a single sampling trajectory: the large per-sample variance masks the overall mean shift and distorts frequency measurements. To address this issue, we propose ENCORE, a training-free, post-hoc correction that reuses the differences between denoising outputs at consecutive steps, which the sampler otherwise discards, and thus requires no additional network evaluation. In a spectral-then-mean (S2M) pipeline, ENCORE first restores the per-band power via magnitude scaling and then realigns the sample centroid in a pre-trained feature space by ridge regression over these differences. The two stages match a second moment per frequency band and a first moment in feature space, after the last step. Because it operates after sampling, ENCORE is plug-and-play and yields further gains on top of sampler-level optimizations. Evaluations across pixel-space, image latent, and audio latent diffusion models demonstrate that ENCORE improves sample quality with negligible overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.