acceptodds
Under review as a conference paper at ICLR 2027

Same Trajectories, Different Credit: Scoring-Rule Dependence in Tool-Fault Recovery Evaluation

Abstract

After a tool call fails and keeps failing, an agent can stop, do the parts of the task the failure did not block, report the failure, or assert the outcome it was asked for. A benchmark’s continuation rule decides which of these its score credits, so a higher pass rate does not say which behaviours an intervention suppressed. ToolMaze ships two such checks and evaluates its single-path category c1 with the stricter one; we build matched scorers differing only in the forbidden set and re-score identical stored trajectories. That one change moves the pass rate by 11 to 47 points in every arm, and lowers the benefit credited to a recovery message by 6.6 and 16.0 points on two models. A pre-registered replication on sixty held-out tasks reproduces the level shift but not that reduction: score levels are rule-dependent throughout, whereas the attenuation does not replicate across task sets. Which later sub-tasks count as blocked is not specification-free: three plausible tests label the same 62 relations differently, and the same tool can assert the failed step’s outcome in one branch and report the failure in another. Measured apart, the behaviours diverge: the pass rate rises by 33 points while structured-field agreement with a fault-free run on independent sub-tasks falls from 79.4% to 39.7%, and final answers report the failure more often. We recommend reporting the continuation rule with the score, and release the replay protocol, the pre-specified adjudication rubric and the labels.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.