Same Agent, Different Conclusion: Why Recovery-Method Evaluation Is Underspecified
Abstract
Tool-using agents recover from execution failures by retrying a call or switching tools, which is what lets one finish in a complex environment. Every comparison of such methods is scored by a rule that either counts the calls or ignores them, and neither case is safe. Where they are counted, a repair that should have rescued a failure turns a correct episode into a wrong one: a tool returns empty, the agent retries, and the extra call breaks the match against the gold answer. The effect is not incidental: a reference agent, the middleware it ships with, and a benchmark's own checker compose into it, and emptying half the tool returns takes it from to while emitting more calls; the identical runs scored with recovery calls excluded show no drop. Where they are not counted, the winner depends instead on how accuracy is traded against tokens: across episodes over nine recovery methods and five benchmarks, four standard ways of making that trade select multiple winners. Both are decidable before an experiment runs: which case applies is stated in the scoring rule, and the weight at which a ranking changes is a function of two columns. A recovery comparison has to state what counts as success, and what success is worth against cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.