acceptodds
Under review as a conference paper at ICLR 2027

Rethinking On-Policy Distillation in Diffusion Models: Phenomenology and Empirical Insights

Abstract

On-policy distillation (OPD) consolidates diffusion-model specialists into a shared student by supervising its own denoising trajectories. Yet its mechanisms remain underexplored: What does OPD transfer? What determines transfer at student-visited states? When is distillation worthwhile relative to direct joint optimization? We study three representative OPD methods on two backbones and thirteen routes through budget-controlled comparisons, held-out behavioral evaluation, frozen-trajectory interventions, temporal supervision and teacher-source experiments, and comparisons with direct joint reinforcement learning. First, OPD consolidates specialist gains selectively: DanceOPD leads three of four equal-time settings, whereas DiffusionOPD leads three of four equal-data settings, with winners changing in 17 of 52 matched route comparisons. In the Z-Image two-interface equal-data setting, held-out image-editing and multi-image scores improve by approximately 3.4% and 5.8% over the exact initialization. Second, transfer depends on timing and teacher source: query-matched disagreement weighting improves held-out image editing and multi-image generation by approximately 10.3% and 2.8% over uniform weighting, yet fixed low-noise supervision leads on ten of thirteen routes. Moreover, a reinforcement-learning teacher that scores 6.1% above its supervised-fine-tuning counterpart produces a student that scores 2.9% lower in the design case. Third, under UnifiedReward, direct joint reinforcement learning improves general text-to-image, general image-editing, and held-out image-editing scores by approximately 5.5%, 6.6%, and 93.3% over initialization. Equal-data DiffusionOPD is 2.6% better than direct joint reinforcement learning on general image editing but 16.9% and 24.1% worse on general text-to-image and held-out image editing, respectively. We hope these findings inform future OPD design and use by the community. All resources will be publicly released, including code, checkpoints and datasets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.