ToolPlan-Bench: Diagnosing Hidden Planning Failures in Long-Horizon Tool-Calling Agents Beyond Success Rate
Abstract
Large language model (LLM) agents are increasingly deployed on long-horizon, tool-calling tasks, yet they are almost universally evaluated by a single black-box outcome: task success rate. We argue that success rate systematically hides planning failures. An agent can reach the correct final state through a flawed plan—taking redundant actions, executing operations out of order, skipping required preconditions, or stopping prematurely on easy instances—and still be scored as a success. We introduce ToolPlan-Bench, a diagnostic benchmark that decomposes agent behaviour into a structured taxonomy of seven planning-failure modes and measures the right-outcome-wrong-plan rate (ROWP): the fraction of successful trajectories that nonetheless carry a planning-failure label. Built on the real tasks, tools, and executable databases of -bench (retail and airline domains), ToolPlan-Bench evaluates six frontier models on 990 live agent trajectories. We find that ROWP is large and statistically robust—the bootstrap 95% lower bound exceeds 0.23 for every model, meaning at least one in four successful trajectories hides a planning flaw. ROWP rises monotonically with task complexity (from 0.43 to 0.57 as the number of required write operations grows), and 60% of trajectories diverge from the gold plan at the very first step. Five independent human annotators validate our rule-based labels (Cohen's between the human majority and the rule labels on the core label; on the structural dimensions underlying ROWP). Finally, we show that a natural intervention—prompting the agent to plan before acting—can significantly change success rate (in both directions across models) yet has no significant effect on ROWP for any model, demonstrating that the planning failures ToolPlan-Bench measures are robust to naive prompting and constitute an open challenge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.