Agents' Overreliance on Unreliable Tools
Abstract
LLM agents use tools to access information and perform computations beyond their parametric knowledge. Existing tool-use benchmarks evaluate whether agents select and call the right tools, assuming that tool returns are reliable. However, tools can return plausible but incorrect outputs. We evaluate 14 models with three tools (web search, an LLM sub-agent, and a code executor), corrupting their returns to examine whether agents overrely on unreliable tools. Agents adopt corrupted returns at high rates, with mean adoption exceeding one third for every tool and reaching 68.0% for web search. Corrupted returns also frequently override correct answers that agents give without tools, and this reliance persists in more complex tasks and in tasks that combine multiple tools. Agents often make more tool calls under corrupted returns, and reasoning traces show them questioning a return and even stating the correct answer. Nevertheless, they still pass the corrupted content to users, rarely warning them of the conflict. We further test interventions spanning prompts, tool metadata, post-training, and activation steering. Verification prompts reduce adoption across all three tools while largely preserving accuracy with correct returns, and reliability labels provide partial mitigation in web search. The tested post-training methods offer limited gains, and the effect of activation steering depends on the model and intervention strength. Overall, our findings provide a basis for diagnosing tool overreliance and for developing agents that assess tool returns rather than assuming their reliability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.