PACPO: Polarity-Aware Contribution Policy Optimization for Multi-Reward RL
Abstract
As the deployment of large language models expands to complex reasoning, tool-use, and interactive applications, post-training increasingly involves multiple behavioral objectives. To jointly optimize these objectives, multi-reward reinforcement learning has become a key paradigm in language model post-training. Existing methods typically preserve reward-wise information during advantage estimation but merge these signals before applying probability-ratio clipping, causing all reward dimensions to share the boundary response determined by the aggregated advantage. When the policy ratio falls outside the clipping interval, the shared clipping response may mask corrective gradients from reward contributions of the opposite polarity. To address this issue, we propose Polarity-Aware Contribution Policy Optimization (PACPO), which separates reward contributions by polarity and applies polarity-specific clipping responses. This design allows corrective signals from contributions of opposite polarity to remain active, rather than being suppressed by clipping after aggregation. By preserving these otherwise masked signals, PACPO mitigates clipping-boundary gradient masking, thereby promoting more balanced joint optimization across multi-reward objectives. Experiments across diverse multi-reward post-training settings show that PACPO consistently outperforms existing multi-reward optimization methods, demonstrating its effectiveness for multi-reward policy optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.