Learning Where and What to Optimize: From Diffusion Distillation to Reinforcement Learning
Abstract
Teacher-guided distillation provides dense supervision and fast optimization, whereas reward-driven reinforcement learning (RL) directly optimizes scalar rewards at a substantially higher sampling cost. We seek to unify these complementary strengths within a single diffusion post-training framework. A natural alternative is sequential training: distill first, then switch to reward optimization. However, such stage-wise training can suffer from post-switch degradation, late-stage collapse, and sensitivity to the handoff schedule, particularly under limited prompt diversity. We propose a Unified Reward-adaptive Framework (URF) that jointly determines where to optimize and what velocity target to follow for each generated sample. Both decisions are derived from a single reward-adaptive objective and controlled by the same sample-wise reward-preference weight p ∈ [0, 1]. For the velocity target, motivated by the coarse negative-sample update directions, URF uses negative reward feedback to constrain teacher-guided correction rather than directly applying the repulsive NFT update. As the reward increases, p smoothly transitions the update from DiffusionOPD to DiffusionNFT, eliminating the need for an explicit stage-wise switch. To our knowledge, URF is the first diffusion post-training framework to unify teacher-guided distillation and reward-driven RL at both the query and velocity-target levels within a single reward-adaptive objective. Extensive experiments show that URF combines substantially faster convergence than teacher-guided distillation with continued reward-driven improvement, achieving 0.963 on both GenEval and OCR while converging 1.7–7.9× faster than DiffusionOPD in multi-reward training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.