FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification
Abstract
Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing density-approximated policy gradient methods rely on stochastic SDE samplers to construct tractable transition kernels, which introduce training-inference inconsistencies and necessitates Classifier-Free Guidance (CFG). DiffusionNFT bypasses trajectory likelihoods through implicit velocity regression, but uses fixed residual coefficients in its branch targets, with rewards controlling their relative loss weights. We propose Flow Advantage-Weighted Rectification (FlowAWR), a paradigm that recasts continuous generative policy optimization as supervised regression toward a theoretically optimal velocity field. Starting from the optimal policy of a KL-constrained reward maximization, FlowAWR derives the optimal velocity field that admits a magnitude-aware, advantage-weighted rectification form, yielding SDE-free optimization. In comparative evaluations on SD3.5-Medium, FlowAWR achieves improved alignment performance alongside a 2 to 5 convergence acceleration over DiffusionNFT (e.g., reaching a 24.12 PickScore in 1.2k steps, versus 23.82 in 2.0k steps for DiffusionNFT and 23.50 in 4k steps for FlowGRPO). Under multi-reward constraints, FlowAWR sustains quality, satisfying structural rules with stable out-of-domain performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.