acceptodds
Under review as a conference paper at ICLR 2027

RISE-DMD: Reward-Improved Self-Evolving Distribution Matching Distillation for Black-Box Reward Optimization

Abstract

Distribution Matching Distillation (DMD) can distill a multi-step diffusion model into a deterministic few-step generator, but introducing black-box rewards into DMD poses a key challenge: black-box rewards provide only scalar evaluations of final samples and do not directly specify the target distribution or marginal velocity field required by DMD. We propose RISE-DMD (Reward-Improved Self-Evolving Distribution Matching Distillation) and introduce Reward Score to learn a reward-improved marginal velocity field. Reward Score is trained on on-policy samples generated by the current student and their reward feedback, after which DMD uses the learned field to update the few-step student. The updated student then generates data for the next training round, allowing the target distribution to evolve together with the student distribution. Theoretically, we characterize the population-optimal Reward Score field and the corresponding reward-tilted target distributions at each diffusion time, and provide a regularized reward-optimization interpretation under a fixed-reference constraint. In the main SD3.5 Medium experiment, RISE-DMD outperforms the compared methods on all six preference and quality metrics, three of which are not used for training. Results on Sana 1.6B, DrawBench, and OCR further demonstrate the method's applicability across model backbones and prompt distributions, as well as its ability to optimize non-differentiable task rewards using only scalar feedback.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.