SAW: Stage-Aware Weighting for Multi-Reward Reinforcement Learning
Abstract
Multi-reward reinforcement learning for language models combines objectives that can improve at different rates. Fixed aggregation does not adapt their weights as training progresses, while gradient-based weighting introduces additional gradient computation. We propose Stage-Aware Weighting (SAW), a lightweight method that adjusts objective weights using coefficients of variation computed from the current rollout batch. SAW operates at the reward level in GRPO and the advantage level in GDPO. Its weighting rule requires neither additional backward passes nor statistics from previous batches. We evaluate SAW on tool calling and summarization. Across six BFCL-v4 settings, SAW improves overall accuracy over default fixed aggregation by 1.09–7.37 percentage points and achieves higher Overall point estimates than the evaluated gradient-based weighting baseline in all six settings. In the 1.5B two-reward setting, SAW also exceeds all three tested fixed ratios in Overall. Summarization gains persist under two evaluation judges not used for training. These results support current-batch reward statistics as a practical and computationally lightweight basis for adaptive multi-reward optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.