Rectified Preference Alignment via Deterministic Policy Optimization
Abstract
Preference alignment for text-to-image rectified-flow models commonly relies on image-level rewards. However, such supervision is flow-blind: rectified and non-rectified paths to the same preferred image receive identical rewards. Consequently, alignment can improve endpoint reward without preserving the pretrained flow geometry that supports effective finite-step sampling. This may compromise generation quality and weaken alignment gains as the number of sampling steps decreases. We address this gap with Flow-VPA, a Flow-aware Velocity Preference Alignment framework that combines rectified-transport supervision with deterministic policy optimization. To provide this supervision, we introduce BT-Flow, a Bradley–Terry reward model over state–velocity pairs that lifts existing chosen–rejected image comparisons into local velocity preferences. Conditioned on a shared intermediate state, time, and prompt, BT-Flow learns to favor the velocity defining a straight-line completion to the chosen image over the corresponding velocity toward the rejected image. Building on this reward, Flow-VPA treats predicted velocities as deterministic actions and updates the policy using reward gradients at on-policy states. This enables direct alignment of the deterministic flow without introducing an auxiliary stochastic differential equation, evaluating trajectory likelihoods, or differentiating through the sampler. Experiments on SD3.5-M show that Flow-VPA improves preference alignment and generation quality over representative baselines under practical finite-step sampling, with a better trade-off between alignment performance and policy-training compute.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.