acceptodds
Under review as a conference paper at ICLR 2027

Tool Graph Complexity for LLM Tool-Calling: A Reproducible Benchmark Auditing and Generation Framework

Abstract

Language-model agents are increasingly evaluated by how many tools a task requires: benchmarks stratify difficulty by tool count, and leaderboards report accuracy falling as tool lists grow. We argue that this number conflates two different things. One is a property of the benchmark: the reference tool set an agent is scored against often lists tools the task does not need, so recall penalizes any agent that finds a shorter correct plan. The other is a property of the task: how the required tools must compose. We introduce AgentGraphBench, a framework that separates them. It builds a dependency graph over a tool catalog directly from tool schemas, audits an existing benchmark's reference sets against an executable runtime and human-anchored judges, and generates new tasks whose reference set is a sampled subgraph of chosen topology. On StableToolBench, a third to a half of reference tools turn out to be unnecessary, and scoring against the necessary ones raises every agent's recall by about a quarter: much of the reported count-driven degradation is an annotation artifact. On generated tasks with clean references, difficulty is dominated by subgraph size, but at equal size sequential composition stays harder than fan-out or fan-in, and the penalty disappears once dependencies are written out explicitly. What makes tool use hard, then, is inferring dependencies that tool descriptions leave implicit, and the tools agents omit are precisely those the instruction fails to motivate, not those the runtime can do without. We release the framework, a cleaned StableToolBench subset and 2,180 topology-controlled tasks, and we map where the structural ranking transfers: to execution and to a second catalog only in coarse form.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.