acceptodds
Under review as a conference paper at ICLR 2027

PACPO: Polarity-Aware Contribution Policy Optimization for Multi-Reward RL

Abstract

As the deployment of large language models expands to complex reasoning, tool-use, and interactive applications, post-training increasingly involves multiple behavioral objectives. To jointly optimize these objectives, multi-reward reinforcement learning has become a key paradigm in language model post-training. Existing methods typically preserve reward-wise information during advantage estimation but merge these signals before applying probability-ratio clipping, causing all reward dimensions to share the boundary response determined by the aggregated advantage. When the policy ratio falls outside the clipping interval, the shared clipping response may mask corrective gradients from reward contributions of the opposite polarity. To address this issue, we propose Polarity-Aware Contribution Policy Optimization (PACPO), which separates reward contributions by polarity and applies polarity-specific clipping responses. This design allows corrective signals from contributions of opposite polarity to remain active, rather than being suppressed by clipping after aggregation. By preserving these otherwise masked signals, PACPO mitigates clipping-boundary gradient masking, thereby promoting more balanced joint optimization across multi-reward objectives. Experiments across diverse multi-reward post-training settings show that PACPO consistently outperforms existing multi-reward optimization methods, demonstrating its effectiveness for multi-reward policy optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.