acceptodds
Under review as a conference paper at ICLR 2027

Same Agent, Different Conclusion: Why Recovery-Method Evaluation Is Underspecified

Abstract

Tool-using agents recover from execution failures by retrying a call or switching tools, which is what lets one finish in a complex environment. Every comparison of such methods is scored by a rule that either counts the calls or ignores them, and neither case is safe. Where they are counted, a repair that should have rescued a failure turns a correct episode into a wrong one: a tool returns empty, the agent retries, and the extra call breaks the match against the gold answer. The effect is not incidental: a reference agent, the middleware it ships with, and a benchmark's own checker compose into it, and emptying half the tool returns takes it from to while emitting more calls; the identical runs scored with recovery calls excluded show no drop. Where they are not counted, the winner depends instead on how accuracy is traded against tokens: across episodes over nine recovery methods and five benchmarks, four standard ways of making that trade select multiple winners. Both are decidable before an experiment runs: which case applies is stated in the scoring rule, and the weight at which a ranking changes is a function of two columns. A recovery comparison has to state what counts as success, and what success is worth against cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.