acceptodds
Under review as a conference paper at ICLR 2027

FACTOR-OPD: Allocate Reward-Specific Credit Before Combining Policy Gradients

Abstract

Multi-reward reinforcement learning must balance objectives that can require different updates within the same response. Sequence-level reward aggregation leaves this token-credit problem unresolved, while teacher feedback from uncontrolled references can mix the effects of several reward properties. We introduce FactorOPD, which allocates reward-specific token credit before combining objective gradients. Verifier-checked factorial references contrast high and low levels of each reward property while balancing the other verified factors. A frozen teacher scores identical student-sampled tokens under these references, and constrained allocation converts the resulting contrasts into positive, bounded, mean-one token weights. This preserves each component advantage's sign and response-wise mean. Parameter-space MGDA then combines the induced objective gradients, accounting for their geometry in the shared model parameters. Across six benchmarks spanning mathematical reasoning, code generation, and instruction following, FactorOPD outperforms GRPO, GDPO, and GD²PO with both Qwen3-4B and Qwen3-8B. Domain-average accuracy gains over GD²PO range from 1.73 to 1.84 percentage points. Ablations support complementary benefits from token allocation and gradient coordination, including the contribution of token alignment. Reward-wise analyses further show higher correctness without reduced format compliance on Math and Code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.