Testing Agent Process Verifiers with Counterfactual Instructions
Abstract
Process verifiers guide tool-using agents by ranking candidate actions. However, evaluating a verifier on a single instruction cannot tell whether its preference actually depends on the task requirement. We introduce crossed instruction-trajectory matching: the same two candidate trajectories are evaluated under two instructions that differ in one requirement. A separate action-level predicate recomputes the required argument from the instruction and shared history, and verifies which candidate action is correct. A verifier must reverse its preference to solve the group, ruling out a fixed candidate preference. Using this design, we construct 195 groups from three tool-use benchmarks and evaluate five prompted judges and five trained reward models. On selection-by-attribute groups, 12 of 42 judge-group pairs answered correctly under the original instruction fail under the edited instruction, including seven fixed candidate preferences. Judges with similar crossed accuracy fail in different ways, including fixed preferences, position-following, and ties. Organizing the listed options by item improves all five judges on both diagnosed and held-out groups, whereas simply restating the criterion does not consistently help. Raw-value tables, which group observed values by item without computing attributes, recover 5-10 of 11 held-out groups, and computed tables recover 8-11. Because success requires correct choices under both instructions, this recovery cannot be explained by a stronger fixed preference for one candidate. Finally, we show that evaluation protocol matters: scoring candidates separately can lose distinctions recovered by direct comparison, and aggregating step-level scores by their minimum can erase correct local error detection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.