TGPO: Tracked-Gradient Policy Optimization for Multi-Reward Reinforcement Learning
Abstract
Multi-reward reinforcement learning requires coordinating objectives whose signals differ in scale, variability, and direction. GDPO preserves reward distinctions through decoupled normalization, but reward statistics alone do not determine gradient compatibility. Sampling noise further affects nonlinear coordination across minibatches. We introduce Tracked-Gradient Policy Optimization (TGPO), which retains decoupled normalization and coordinates temporally tracked objective gradients. A projected recursive estimator maintains a bounded state for each objective. TGPO approximately maximizes the minimum predicted surrogate improvement within a relative neighborhood of the mean ascent direction. The same tracked vectors define both the interaction geometry and the gradient supplied to the optimizer, making temporal information part of the update to support stable gradient coordination. We evaluate TGPO against GDPO and DVAO on tool use, mathematical reasoning, and coding reasoning. Evaluation covers correctness metrics and constraint adherence. Compared with these baselines, TGPO achieves state-of-the-art aggregate performance across the evaluated settings, demonstrating the effectiveness and broad applicability of temporal gradient coordination in multi-reward reinforcement learning. Our code is available at https://anonymous.4open.science/r/TGPO-072E/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.