One Stage Few Step Distillation with On Policy Trajectory Anchoring
Abstract
Preference-aligned diffusion models achieve strong generation performance but typically require many sampling steps. Developing few-step diffusion models that retain strong alignment with human preferences remains an important research goal. On-policy distillation (OPD) transfers teacher-aligned capabilities through local supervision, but its performance degrades sharply with fewer sampling steps. Distribution Matching Distillation (DMD) enables few-step generation but can introduce distributional drift and weaken aligned capabilities. Separating the transfer of teacher-aligned capabilities and few-step distillation into two stages can cause the second stage to undermine the aligned capabilities or few-step generation ability learned in the first. To address these limitations, we propose On-Policy Distribution Matching Distillation (OPDMD), a single-stage framework that jointly enforces teacher anchoring and distribution matching along student-generated trajectories. OPDMD anchors the student velocity to the teacher at each selected student state and applies a distribution-matching correction at the corresponding successor state. We unify these two supervision signals into a transition-aware MSE objective that scales the correction by the transition span. Experiments on SD-3.5-Medium with GRPO-aligned expert teachers show that OPDMD better preserves aligned capabilities and distributional fidelity than OPD, DMD, and their sequential combinations. Additional analyses show that transition-aware correction is especially beneficial under sparse sampling and that balancing the two supervision signals is important for overall performance. These results establish OPDMD as an effective approach to preference-preserving few-step generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.