Generalizing On-Policy Self-Distillation for Diffusion Models
Abstract
Scalar image rewards can guide diffusion model training, but translating them into effective parameter updates remains challenging. DiffusionOPSD offers an explicit target-fitting route, yet reward-gradient-based target construction still requires differentiable evaluators. Thus, we present GenOPSD, an on-policy self-distillation method for score-only feedback. It converts scores of decoded images from completed original trajectories and projected self-SDE branches into detached query-level targets, then fits the clean-latent targets with uniform positive clean-prediction MSE and refreshes the policy. Rewards affect target construction only, without gradients through the evaluator or decoder or a reward-weighted policy-gradient loss. Matched single-reward SD3.5-M results support the generalization: GenOPSD approaches DiffusionOPSD on four differentiable reward metrics, staying within 5.6% across all four and gaining 1.5% on ImageReward, while also supporting score-only GenEval and OCR evaluation. Under multi-reward training, its scores remain within 7.8% of DiffusionNFT across measures. On Z-Image-Turbo, it improves over the base by 2.5% on GenEval and 1.2% on OCR. On Qwen-Image-Bench, it improves over the SD3.5-M base by 34.2% overall and by 8.9% to 62.5% across five categories. Ablations support the roles of both state sources, prompt-population-relative signed preferences, coverage over native noise levels, and policy refresh. Guidance tests favor the CFG 1 actor, and qualitative examples show stronger composition and text rendering but residual color and season errors. We hope GenOPSD offers the community an effective method for training diffusion models with scalar reward feedback.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.