Reinforcing Diffusion Models via Value-Implicit Reward Matching
Abstract
Online Reinforcement Learning (RL) has proven effective for aligning diffusion and flow models with task-specific rewards. However, existing methods rely heavily on newly generated on-policy samples for stable policy improvement, leading to poor sample efficiency. We introduce **VIRM**, a value-implicit reward matching framework that casts policy improvement as regression on replayed samples. VIRM builds on the optimality condition of KL-regularized RL, where the optimal policy exponentially tilts the reference toward high reward. Inverting this identity reveals that each sample's reward is implicitly encoded by the policy–reference log-density ratio and a per-prompt log-partition term. Using an exponential moving average (EMA) policy as the reference, VIRM optimizes the current policy by matching this policy-implied reward to external feedback. Concretely, the log-density ratio is estimated via an ELBO difference that acts as an implicit value, while the intractable log-partition term is predicted by a token-wise value head. VIRM learns from replayed data without importance weights or stored denoising trajectories. To focus replay on correctable errors, samples are prioritized by reducible loss, a reference-calibrated measure of learnability rather than intrinsic difficulty. On GenEval2, VIRM improves FLUX.2-klein-base-9B from to , outperforming the strongest baseline by more than 40%, with concurrent gains in CLIP score, PickScore, and HPSv2. These results establish value-implicit reward matching as a sample-efficient post-training method for diffusion and flow models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.