Auditing Reward-Model Selection on Tool Trajectories: Identification, Policy Utility, and Natural Transfer
Abstract
Reward models often turn a candidate pool into one selected trajectory. The maximum score by itself does not show whether that trajectory is useful. We audit selected checkpoints on executable tool trajectories. The main evidence comes from 37 mixed-outcome -bench airline pools, with a secondary closure on AgentSuite airline and held-out retail cells. Five independently started fixed-process runs gave identical results on the original 370 candidates. By contrast, three semantically equivalent serializations agreed on the winner in only of pools (95% CI ). At the primary joint-evaluator seed, mixed-only excess utility was , , and in the three cells and varied across seeds. On the same 69 completed mixed pools, the follow-up sensitivity checks found full-cycle winner agreements after reversal of the primary orientation and agreements between padded and dynamic full cycles. A secondary matched-slot arm covered 69/102 gates and provides a feasibility-conditional, binary-outcome, policy-plus-graph diagnostic. Our conclusions are limited to the named checkpoints, policies, pools, serializations, graph constructions, and context protocols; they do not rank selector classes or establish general equivalence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.