acceptodds
Under review as a conference paper at ICLR 2027

ReFormFlow: Reforming Flow Reinforcement via Directly Matching Reward Distributions

Abstract

Online reinforcement learning for flow matching is challenged by costly terminal-density evaluation. Recent approaches introduce stochastic transitions for likelihood-ratio updates or use local reward-weighted velocity regression. This motivates a terminal-sample objective that retains deterministic generation while specifying the reward-tilted distribution toward which the policy is optimized. We propose ReFormFlow, an on-policy framework motivated by reverse Kullback–Leibler minimization toward an exponentially reward-tilted reference distribution. An auxiliary path-space analysis motivates a velocity-regression surrogate for the terminal log-density ratio. ReFormFlow aggregates self-normalized regression contributions across re-noised timesteps into a single residual against the scaled, group-centered reward, jointly learning the velocity field and a prompt-conditioned log-partition. Unlike per-step regression, reward supervision applies to the combined contribution. Both training rollouts and inference retain deterministic ODE sampling without classifier-free guidance. Experiments cover single- and multi-reward optimization and out-of-domain evaluation. ReFormFlow is up to faster than DiffusionNFT at matched reward levels in the single-reward comparisons, measured in sampling-and-training GPU-hours. On PickScore, it achieves 24.20 in about 1k updates, compared with 23.82 in 2k updates for DiffusionNFT.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.