FlowSRL: Stable and Scalable Reinforcement Learning for Diffusion and Flow Models
Abstract
Online reinforcement learning has established itself as the standard approach for post-training diffusion and flow-based generative models with arbitrary reward signals. Despite their limited success, they often rely on various training tricks to remain stable, and inevitably collapse the diversity of the generated samples. In this work, we revisit both problems. To address the instability, we introduce deviation clipping, a simple trust region on the velocity update that stabilizes training. Anchoring the update to a learned flow-matching baseline rather than to the policy itself stabilizes it further and speeds up convergence. The resulting algorithm, FlowSRL, matches or outperforms prior diffusion RL methods on standard text-to-image benchmarks at significantly smaller batch sizes and roughly 2× faster convergence. To address the diversity collapse, we study two complementary mechanisms, a passive one that restricts the policy to refining trajectories the pretrained model begins, and an active entropy reward bonus that incentivizes ex- ploration beyond the base model’s distribution. Finally, we apply FlowSRL to hard reward problem and show that, combined with search during the rollout stage, it is able to solve them reliably where the existing baselines without search struggled.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.