Simple Endpoint Reward Finetuning
Abstract
Reward fine-tuning adapts pretrained diffusion and flow models to desired objectives by optimizing a specified reward. For differentiable rewards, transporting reward gradients backwards along trajectories enables gradient computation at noisy states. However, implementing this in practice is prohibitively expensive; existing first order methods imitating this must make aggressive approximations, often estimating the trajectory through few-step samples and evaluating reward gradients off-policy. This motivates the development of a method using only reward gradients at clean samples, while still approximating the true gradient direction. To this end, we introduce SIMPLE ENDPOINT REWARD FINETUNING (SERF), a remarkably simple fine-tuning objective that avoids gradient transport by directly using the reward gradient at the endpoint. Specifically, at renoised states the trained velocity is regressed against the reference velocity shifted by the clean image's reward gradient. Although easy to state and implement, we show that SERF has a rigorous derivation through stochastic optimal control, and that our objective provides a consistent estimate of the true reward gradient direction. When compared with existing methods, SERF achieves the most favorable trade-offs when used to post-train Stable Diffusion 3.5 on PickScore and HPSv2.1, demonstrating its ability to learn rewards while maintaining diverse outputs. SERF is shown to be highly efficient, up to 10 times faster than Adjoint Matching, while remaining flexible enough to allow endpoint sampling in as few as 5 steps.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.