acceptodds
Under review as a conference paper at ICLR 2027

Distribution-based Regularization for reward fine-tuning diffusion and flow models

Abstract

Reward fine-tuning is an effective way to adapt pretrained diffusion and flow-matching models to downstream objectives, but aggressive optimization often leads to mode collapse and distributional drift. Existing methods attempt to control this drift with reverse-KL regularization, implemented as an L2 penalty on the velocity field; however, this functions by indirectly constraining sampling dynamics rather than directly acting on the generated distribution. To address this, we propose DFR (Distributional Functional Rewards), which augments the task reward with a distribution-dependent, per-sample reward that acts explicitly on the generated distribution. Formally, we generalize the KL-regularized objective from expected rewards, which are linear in the model's distribution, to convex functionals of it. The first variation of this objective defines an effective, distribution-dependent reward that directly penalizes deviation from the pretrained distribution. This formulation yields additive per-sample rewards for the Fréchet distance (FD), maximum mean discrepancy (MMD), and DreamSim diversity, recovering several recently proposed rewards as special cases. Across various text-to-image fine-tuning tasks, DFR matches the peak reward of KL-tuned baselines while increasing diversity by 11-29% at equal or lower drift.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.