Rethinking On-Policy Distillation in Diffusion Models: Phenomenology and Empirical Insights
Abstract
On-policy distillation (OPD) consolidates diffusion-model specialists into a shared student by supervising its own denoising trajectories. Yet its mechanisms remain underexplored: What does OPD transfer? What determines transfer at student-visited states? When is distillation worthwhile relative to direct joint optimization? We study three representative OPD methods on two backbones and thirteen routes through budget-controlled comparisons, held-out behavioral evaluation, frozen-trajectory interventions, temporal supervision and teacher-source experiments, and comparisons with direct joint reinforcement learning. First, OPD consolidates specialist gains selectively: DanceOPD leads three of four equal-time settings, whereas DiffusionOPD leads three of four equal-data settings, with winners changing in 17 of 52 matched route comparisons. In the Z-Image two-interface equal-data setting, held-out image-editing and multi-image scores improve by approximately 3.4% and 5.8% over the exact initialization. Second, transfer depends on timing and teacher source: query-matched disagreement weighting improves held-out image editing and multi-image generation by approximately 10.3% and 2.8% over uniform weighting, yet fixed low-noise supervision leads on ten of thirteen routes. Moreover, a reinforcement-learning teacher that scores 6.1% above its supervised-fine-tuning counterpart produces a student that scores 2.9% lower in the design case. Third, under UnifiedReward, direct joint reinforcement learning improves general text-to-image, general image-editing, and held-out image-editing scores by approximately 5.5%, 6.6%, and 93.3% over initialization. Equal-data DiffusionOPD is 2.6% better than direct joint reinforcement learning on general image editing but 16.9% and 24.1% worse on general text-to-image and held-out image editing, respectively. We hope these findings inform future OPD design and use by the community. All resources will be publicly released, including code, checkpoints and datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.