acceptodds
Under review as a conference paper at ICLR 2027

SWIFT: Self-guided Weak-to-strong dIFfusion models via Target extrapolation

Abstract

Diffusion models are typically trained to predict clean samples, while autoguidance improves generation at inference by extrapolating a trained denoiser away from a weaker one. We introduce SWIFT, which moves this weak-to-strong extrap- olation into training: each clean-sample regression target is extrapolated away from the prediction of a frozen weak denoiser, which is discarded after training. Although this deliberately shifts the target away from the posterior mean, a spectral analysis of finite-time population gradient flow shows that a nonempty range of extrapolation strengths can yield a denoiser strictly closer to the original posterior mean than matched standard training. Experiments on CIFAR-10 and ImageNet- 256 show consistent improvements across model scales and latent representations. On ImageNet-256, SWIFT improves unguided FID across every evaluated con- figuration, and the advantage persists with longer training. SWIFT also remains complementary to inference-time autoguidance and classifier free guidance: their combination gives the best FID across all tested configurations. On the DINOv3-L RAEv2 configuration, SWIFT reaches 1.15 gFID without inference-time guidance and 0.99 with retuned guidance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.