Multi-Reward Hacking: From Aggregation-Induced Drift to Trust-Anchored Defense
Abstract
Modern LLM post-training increasingly combines several reward signals into one training objective. We show that symmetric aggregation of such signals can induce multi-reward hacking: the aggregate reward improves while performance on a primary criterion degrades, as the policy increasingly adopts a behavior that satisfies an auxiliary channel without satisfying the primary one. Formulating post-training as a one-step multi-reward MDP, we analyze a two-channel, three-action instance and prove that, under symmetric aggregation, the deterministic conditional-mean GRPO dynamics select the auxiliary action whenever the primary action is rewarded less often outside a joint-success component that both actions share; equal access to joint success does not remove the drift that this reward-frequency gap induces. Whether this selection is harmful depends on the unknown true reward, which we keep separate from the channels. The failure raises two questions: which channel should receive priority and how. Trust-Anchored Gating (TAG) first elects an anchor channel from the drift that each channel induces when optimized alone, then conditions auxiliary credit on the elected anchor structurally; static reweighting toward the anchor is instead calibration-sensitive, and every fixed weight we test still collapses. In the same model TAG selects the primary action from every initialization whenever the objectives are not always compatible, for every finite gating weight. Across Qwen and Llama models on mathematics and program synthesis, symmetric, constrained and regularized aggregators show the predicted degradation while TAG keeps primary performance while retaining useful auxiliary gains when the rewards are compatible.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.