Compile the Catalog, Not the Task: ToolGraph as a Tool-Use Agent Harness that Updates, Scales, and Transfers
Abstract
Tool-use agents must chain interdependent calls, and which outputs satisfy which later inputs is a property of the deployed tool catalog, not of any one task or model. We argue that this structure should be compiled once into a persistent artifact that any frozen agent can read. ToolGraph is a directed, weighted tool-dependency graph built from tool interfaces alone: parameter matching and counterfactual probes, in which an evaluator LLM compares a candidate chain with the same query answered without its proposed predecessor, propose edges, and a Bayesian combination of match and information-gain evidence sets each edge's confidence. Construction reads no task data, reward, or expert label. A frozen planner reads the graph through a top- injection ranked by personalized PageRank or by a goal-conditioned action value composed from graph paths. We study ToolGraph as an agent harness along three axes. Update (continual learning): the agent stays frozen while the graph learns. Create/read/update/delete (CRUD) operations driven by execution outcomes raise completion on held-out -Bench tasks by – points for two planners, with the gain saturating by – of the update stream. Scale (test-time scaling): spending test-time compute on an LLM scorer that reads graph evidence outperforms Language Agent Tree Search in all six ToolComp and -Bench settings while using – fewer LLM calls per task. Transfer: with no training, subgraphs retrieved from a -node ToolGraph compiled on Toolathlon's catalog lift DeepSeek-V3.2 on -Bench from to , while the reverse direction trails the target's own graph. Across three benchmarks and three backbones, the reader matters more than the graph: the same edges help or hurt depending on how the planner reads them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.