RewardFlow: Learning to Align Flow Models with Reward Distributions
Abstract
Post-training alignment is important for flow-based text-to-image generation, but existing GRPO-style methods usually optimize scalar rewards and tend to concentrate the policy on a few reward-favored samples. This limits diversity, weakens generalization to unseen preference metrics, and becomes unstable under multi-reward optimization. In this paper, we propose RewardFlow, a reward distribution matching framework for post-training flow matching models. Instead of directly maximizing scalar rewards, RewardFlow constructs a reward-induced distribution over sampled denoising trajectories and aligns the flow policy with this distribution, allowing multiple high-quality generations to receive positive optimization signals. To stabilize image flow optimization, we use a warm-started learnable prompt-conditioned partition function for reward calibration and a step-wise KL constraint to preserve the pretrained flow prior. Experiments on SD3.5-M and FLUX.1-dev under both single-reward and multi-reward settings show that RewardFlow achieves stronger preference alignment, better generalization, and more diverse generations. Notably, under multi-reward training on SD3.5-M, RewardFlow improves HPS-v2.1 from 0.326 to 0.362 and the unseen Aesthetic Score from 6.226 to 6.470 over Flow-GRPO. Our code will be released on GitHub.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.