StepX-OPD: Rethinking Multi-Teacher On-Policy Distillation in Diffusion Models
Abstract
Text-to-image diffusion models must combine broad capabilities, such as instruction following, text rendering, perceptual quality, and human preference, with the few-step sampling that deployment requires. On-policy distillation supervises a student on states it generates itself, consolidating capabilities from several specialized teachers into one model, while step distillation compresses a sampling trajectory into a few steps; we study combining the two. However, existing pipelines instead chain them: a full-step multi-teacher model is trained first and compressed only afterward, so compression waits on that stage, and the capabilities it consolidates are not guaranteed to survive compression. To address this, we introduce StepX-OPD, a multi-teacher on-policy distillation method that jointly consolidates specialist capabilities and compresses the sampling trajectory: after a brief self-distillation warmup, each mini-batch is drawn from one training source and routed to the frozen reward-specialized teacher for that source, which drives distribution matching on states supplied by the student's own few-step rollout. Furthermore, on low-noise batches the routed teacher also regresses the student's velocity field directly, balancing capability transfer against the fidelity of the few-step trajectory itself. Baseline evaluations span SD3.5-Medium and FLUX.1-dev, scoring GenEval, OCR, DeQA, PickScore, and their normalized average. The quantitative result reported here for StepX-OPD is on SD3.5-Medium, where it attains the highest Average among the four-step methods compared and exceeds the Average of the forty-step Flow-OPD model. Jointly optimizing capability integration and step compression on one student trajectory avoids using a consolidated full-step model as the teacher for a later compression stage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.