acceptodds
Under review as a conference paper at ICLR 2027

Invocation-Level Reliability of Tool-Using Agents

Abstract

An early tool-call error can corrupt every downstream input even when an agent subsequently applies the right local rule. We isolate this effect with matched evaluations under a correct baseline history and the model's own free-running history. For tasks of depth , their pooled invocation rates and define a fit-free relative propagation loss . Across five open-weight models on controlled routing tasks, the two 7-8B models lose about 68% of baseline capability by depth 6; stronger models mostly remain at ceiling. Our main result is a construct-validity condition for trajectory scoring. Let be the canonical next target after a divergence and the agent-observable history. If has conditional guessing probability at most given , then any fixed-gold scorer credits the agent with probability at most , regardless of its local competence. Gold-agreement severity is therefore driven to a scorer-imposed boundary, and observed reconvergence cannot measure behavioral recovery. Fixed reference alone is not sufficient: the result applies when the canonical target becomes unreachable or unpredictable, and excludes reconstructible targets and evaluators that accept alternative states or milestone paths. Our hidden-constant tasks instantiate the condition: 0 of 869 corrupted-context calls match gold, and 0 of 580 opportunities reconverge. Conditional-on-state replay instead asks whether a call correctly continues from the state the agent actually holds; it yields interior severity estimates of 0.149 and 0.316 and requires no new model queries. Alternative absolute-gap and error-amplification summaries give the same qualitative depth trend.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.