acceptodds
Under review as a conference paper at ICLR 2027

Concordance-Guided Rollout Reweighting and Resampling for Multi-Reward Image Generation Post-Training

Abstract

Multi-reward post-training of text-to-image models commonly reduces several rewards to a single scalar reward. However, the resulting scalarized advantage can disagree in sign with single-reward advantages, concealing conflicting supervision within a rollout group generated for the same prompt. To address this issue, we introduce Scalar Reward Concordance (SRC), a continuous score that quantifies the weakest active-reward support for each rollout's scalarized advantage under a fixed reward scalarization. Derived from within-rollout gradient alignment for the GRPO family, SRC uses only observed reward values. It identifies concordant positive and negative supervision, including rollouts with jointly high or jointly low rewards. Building on this measure, we develop two methods: SRC-Reweight redistributes training weights over existing rollouts without additional sampling, while SRC-Evolve draws additional on-policy rollouts and selects a fixed-size population using SRC. Both methods integrate with GRPO-Guard and DiffusionNFT while preserving the prescribed reward scalarization. We further introduce Joint Success Rate (JSR) to compare models in terms of their ability to satisfy multiple reward requirements simultaneously. Experiments covering five training rewards and three held-out evaluation rewards show improved aesthetics with prompt following broadly retained. SRC-Evolve also improves the joint success rate on the applicable training rewards relative to baselines with matched numbers of generated rollouts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.