From Preference to Flow: Turning Pairwise Preferences into Local Flow Supervision
Abstract
Preference optimization has emerged as an effective post-training paradigm for aligning visual generative models with human preferences. Prior methods usually use pairwise preferences as ranking signals over final outputs or as objectives over the generation trajectory, but they rarely turn such preferences into explicit local supervision for intermediate flow dynamics. Motivated by this gap, we propose **P2Flow (Preference-to-Flow)**, a preference post-training framework that converts pairwise preferences into local state-velocity supervision for intermediate flow dynamics. Given a pretrained flow-based generator, P2Flow uses the latent displacement from the preferred output to the rejected one as a pair-specific preference-violating direction to construct perturbed intermediate states around the preferred flow states, and pairs them with corrective velocity targets whose implied clean endpoint remains anchored to the preferred latent. A timestep-dependent modulation further creates bridge-shaped state perturbations that vanish at both the clean and Gaussian endpoints. This approach retains the standard squared velocity-regression objective of flow matching and requires no additional network modules, reward models, or online rollouts during optimization. In extensive experiments using our 40.8K high-confidence human-VLM consensus preference pairs, P2Flow consistently outperforms supervised fine-tuning and strong preference-optimization baselines on Wan2.1-T2V-14B and LingBot-Video-Dense across three benchmarks, showing clear improvements in both automatic metrics and human evaluations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.