Stable Signed Flow Reinforcement
Abstract
Reward-based finetuning of flow models requires promoting high-reward samples and suppressing low-reward ones while maintaining training stability. Reward-weighted regression provides a simple starting point, but nonnegative weights give no explicit unfitting signal, and repeated reweighting can concentrate the distribution and cause collapse. We propose Stable Signed Flow Reinforcement (SIFR), which uses signed advantages to explicitly reinforce high-reward samples and unfit low-reward ones, while constraining velocity drift from a reference model to stabilize the updates. Alternating updates of the model and a Lagrange multiplier adapt the regularization strength to a prescribed drift budget. Experiments show that SIFR achieves higher rewards than existing methods in fewer training rounds on SD3.5 Medium (2.5B), FLUX.2 Klein (9B), and Cosmos 3 Super Text2Image (64B), using preference rewards as well as rule-based rewards. These gains further transfer to held-out rewards and prompt sets. Ablations show that signed advantages and the drift constraint each contribute to stable training, and that together they make flow RL faster and raise its reward ceiling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.