acceptodds
Under review as a conference paper at ICLR 2027

On-Policy Distillation for Few-Step Diffusion Models

Abstract

Few-step diffusion models promise interactive generation, but compressing a long teacher trajectory into a few student transitions often degrades fidelity. Existing approaches expose a trade-off. Trajectory and consistency distillation retain paired teacher targets but commonly query states that the few-step student does not visit, while distribution matching follows student-generated samples but replaces those targets with indirect distributional feedback that can favor modes already represented by the student. We present OPTD, which resolves this trade-off by rolling the student to an inference state and matching one coarse transition to the endpoint obtained by integrating a frozen teacher through the same interval. The base method therefore combines on-policy states with dense paired supervision using one trainable network. We add a boundary-matched GAN that compares the student's next-boundary latent with a same-condition real latent at the identical noise level in the frozen teacher's feature space. One endpoint construction supports image, bidirectional video, and causal video generation. Across image and video experiments, OPTD delivers strong and stable few-step generation quality across the evaluated settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.