acceptodds
Under review as a conference paper at ICLR 2027

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning

Abstract

Multi-reward reinforcement learning for large language models combines task accuracy with auxiliary objectives such as length and format compliance. In Group Relative Policy Optimization, Reward Combination normalizes the scalarized reward, while Advantage Combination normalizes each reward separately before applying fixed coefficients. The latter removes differences in reward dispersion before combination, even as those statistics change during training. We propose **D**ynamic **V**ariance-adaptive **A**dvantage **O**ptimization (**DVAO**), which adapts the coefficients of normalized advantages using each reward's empirical standard deviation within a rollout group. We prove that DVAO's pointwise advantage magnitude is bounded by that of Reward Combination and characterize the cross-objective coupling in its sensitivity to reward perturbations. An equivalent groupwise rescaling representation relates this magnitude to reward correlations. Experiments on mathematical reasoning and tool use cover four *Qwen3* and *Qwen2.5* model settings. DVAO achieves the highest average accuracy among the compared multi-reward methods in each setting, with length compliance above 99.9% on both math models and the highest average format compliance on both tool-use models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.