acceptodds
Under review as a conference paper at ICLR 2027

SAW: Stage-Aware Weighting for Multi-Reward Reinforcement Learning

Abstract

Multi-reward reinforcement learning for language models combines objectives that can improve at different rates. Fixed aggregation does not adapt their weights as training progresses, while gradient-based weighting introduces additional gradient computation. We propose Stage-Aware Weighting (SAW), a lightweight method that adjusts objective weights using coefficients of variation computed from the current rollout batch. SAW operates at the reward level in GRPO and the advantage level in GDPO. Its weighting rule requires neither additional backward passes nor statistics from previous batches. We evaluate SAW on tool calling and summarization. Across six BFCL-v4 settings, SAW improves overall accuracy over default fixed aggregation by 1.09–7.37 percentage points and achieves higher Overall point estimates than the evaluated gradient-based weighting baseline in all six settings. In the 1.5B two-reward setting, SAW also exceeds all three tested fixed ratios in Overall. Summarization gains persist under two evaluation judges not used for training. These results support current-batch reward statistics as a practical and computationally lightweight basis for adaptive multi-reward optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.