Personalizing Few-Step Diffusion Distillation via Stepwise Reward Augmentation
Abstract
Personalizing a step-distilled diffusion model requires learning a new subject or style while retaining few-step inference. Reference-conditioned on-policy self-distillation (OPSD) provides teacher supervision along the student’s sampling trajectory, but encourages reference fidelity and image quality only indirectly. We propose Stepwise Reward Augmentation (SRA), which combines this supervision with local targets constructed from normalized reward gradients at intermediate denoising states. Schedule-derived weights allocate reward supervision across states, while aggregate gradient-norm balancing controls its strength relative to distillation. Training uses detached rollout states without backpropagating through sampling transitions. A gradient interpretation and matched controls clarify the roles of these design choices and inform practical recommendations for allocating reward feedback across sampling states and balancing it with teacher supervision. Experiments on object and style personalization across two backbones demonstrate improved reward-related reference similarity and image quality over teacher-only distillation and uniform-weight reward baselines. Additional reward recipes, new personalization cases, and training-cost comparisons further support the method’s applicability and competitive efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.