Intervention Success Is Not Mechanism Identification in Tool-Using Agents
Abstract
Tool-using language model agents are often diagnosed by changing prompts, tools, or available information and then attributing the resulting performance change to a particular failure mechanism. We show that this inference can be misleading. We introduce **Causal Agent Bench (CAB)**, a framework for testing such claims with paired tasks, matched controls, model and tool-interface comparisons, and direct interventions. Across several repository-based evaluations, an instruction that contains no task-specific information can improve accuracy, reduce it, or have little measurable effect. This behavior is not tied to one sentence. Across six matched neutral and procedural prompt pairs, Qwen under one setting drops by 11.5 percentage points while every run still produces a valid final answer. A larger study with 30,240 runs finds substantial procedural-minus-neutral differences in ten of twelve settings, although many of the largest differences occur because the model fails to produce a valid tool action. We also show that a successful repair need not identify the original cause of failure. In one controlled experiment, directly providing the source needed after a tool failure removes the observed accuracy gap. These results show that intervention success is not mechanism identification and that controls must be tested in the model and tool interface in which they are used. Code: https://anonymous.4open.science/r/CAB_ICLR27-309E/README.md
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.