Beyond Scores with Reward-Model Embeddings for Diverse Flow RL
Abstract
Reinforcement learning (RL) improves text-to-image flow models, but scalar rewards assigned independently to individual images overlook visual redundancy within a prompt group. We show that reward-model embeddings provide diversity feedback within each prompt group. We propose using regularized D-optimal design to compute diversity weights from reward-model embeddings and reshape scalar rewards before advantage computation. Our method requires no additional image encoder or encoder forward pass and preserves the underlying sampling and policy-update rules. We evaluate our method on Stable Diffusion 3.5 Medium (SD3.5-M) with Advantage Weighted Matching (AWM) and DiffusionNFT, two RL algorithms for flow models. On DrawBench, our method improves overall quality and within-prompt diversity, including gains on three held-out reward models not used during training. On Pick-a-Pic, our method attains each baseline's highest evaluation reward faster for AWM and faster for DiffusionNFT, measured in GPU hours.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.