acceptodds
Under review as a conference paper at ICLR 2027

Multi-Aspect Reward Balance for Diffusion Alignment

Abstract

Image generation quality is inherently multidimensional, so aligning diffusion models with human preferences requires jointly optimizing multiple reward objectives. The central challenge is to improve a single model across reward dimensions through joint training with less manual reward balancing. Rollout samples can be informative for only a subset of rewards, making it important to preserve their distinct supervision during joint training. Common approaches train separate reward specialists, combine rewards with fixed weights, or use manually designed training curricula. However, these approaches do not adequately balance the influence of different rewards during joint training. We propose MARBLE (Multi-Aspect Reward BaLancE), which preserves reward-specific credit through independent advantage estimates and balances their influence using interactions among normalized reward gradients. Our analysis shows that a nonzero normalized refresh direction can locally reduce the participating reward losses together. Gradient diagnostics also show fewer conflicts between the refresh update and individual reward gradients. An amortized formulation of the DiffusionNFT loss, coupled with exponential moving average (EMA) coefficient smoothing, makes this adaptive coordination practical for joint training. On SD3.5 Medium with five rewards, MARBLE improves all eight evaluation metrics over both simultaneous weighted-sum training and fixed uniform coefficients. It achieves OCR 0.96 and GenEval 0.94 while retaining 0.97× weighted-sum throughput, combining multi-reward gains in one model with near-baseline training speed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.