acceptodds
Under review as a conference paper at ICLR 2027

Smooth Then Balance: Gradient Smoothing and Preference Anchoring for Multi-Objective Diffusion RL

Abstract

Multi-objective diffusion reinforcement learning (RL) fine-tunes a single diffusion model to jointly improve multiple objectives (e.g., image quality, human preference, and text rendering). Gradient-based multi-objective optimization methods such as the Multiple Gradient Descent Algorithm (MGDA) avoid manually designed reward weights by combining the gradients of individual rewards into a common descent direction for all rewards. However, we empirically find two drawbacks of MGDA when it is applied to diffusion RL. First, the reward gradients are highly noisy, which makes the resulting update directions unreliable. Second, MGDA determines the weights solely from the gradient directions and can yield highly imbalanced weights that leave some rewards under-optimized. To address these problems, we propose Smooth Then Balance (STB), which first smooths each reward gradient with an exponential moving average before the weights are computed, and then balances the smoothed gradients with a preference anchor that guarantees every reward a minimum weight. Our theoretical results prove that the error of the STB direction is bounded by the gradient estimation error, and that a small STB update guarantees closeness to Pareto stationarity once this error is small. On five-reward fine-tuning of SD3.5-Medium, STB outperforms MGDA on all eight evaluation metrics and achieves the best results on four of the five training rewards among multi-objective diffusion RL methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.