Raw Evaluator Disagreement Can Misrank Trajectories: Interaction-Targeted Agent RL
Abstract
Raw evaluator disagreement can be an unstable policy-learning target in multi-evaluator agent RL. It mixes evaluator-wide strictness with trajectory-specific evaluator dependence, so additive evaluator shifts can change trajectory-penalty rankings even when the interaction profiles are unchanged. We trace this instability to a row-specific cross term and show that, in a fully crossed block with unclipped scores, unrestricted additive shifts can reverse the raw ordering of any two distinct interaction profiles while leaving their interaction-energy ordering unchanged. Interaction-Suppressing Advantage (ISA) addresses this mismatch, relative to an interaction-targeted objective, by double centering the score block before assigning trajectory-level disagreement penalties. With a fixed common baseline and no clipping, its fixed-block advantage is invariant to evaluator-wide additive offsets. In 40 frozen-policy pilot blocks, raw and interaction-energy rankings disagree on 41.98% of trajectory pairs. Matched Raw-Energy/DC-Energy training under the same adaptive energy construction tests whether this target change affects learned policies: double centering lowers rollout discrepancy by 0.0241 (95% CI [-0.0272,-0.0210]), and a subsequent official-success evaluation of the frozen checkpoints yields a +3.64 percentage-point gain. In a separate 40-task Qwen3-8B study with ten seeds per arm, ISA reduces primary held-out global discrepancy by 27.5% relative to mean-minus-SD while improving held-out reward; with the interaction target fixed, a tuned comparison supports the adaptive penalty rule with a +4.36 percentage-point success gain. These results support a conditional design principle: when trajectory-specific evaluator dependence is the intended regularization target and a reliable fully crossed panel is available, remove shared evaluator shifts before forming trajectory-level disagreement penalties.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.