ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
Abstract
We introduce ToolMATH, a math-grounded diagnostic benchmark for evaluating long-horizon tool use under controllable tool-catalog conditions. Its construction pipeline provides a systematic methodology for building tool-use environments: it converts stepwise solutions from the MATH dataset into reusable tools with descriptions and typed schemas, validates their behavior and usability, and applies human repair and authoring to unresolved cases. The human-authored ToolMATHHard split complements the automatic pipeline by recovering unresolved cases, extending evaluation beyond successful automatic trajectories and mitigating validation-related coverage bias. Gold tools and graded distractors model connected operations, overlapping functionality, and missing capabilities. Beyond final accuracy, ToolMATH emphasizes three evaluation axes: (1) Adaptability measures how much Gold-only success is retained when gold tools are replaced entirely by distractors; (2) Robustness measures success retention when distractors are added alongside gold tools; and (3) Tool Connectivity measures answer accuracy as a function of the maximum depth of verified tool-call dependencies. Behavior-conditioned metrics and trace-level failure analyses distinguish reliable tool use, tool avoidance, adaptive substitution, and failures under unreliable catalogs. Together, this construction methodology and these diagnostics provide a controlled testbed for studying how models adapt to changing tool availability, remain robust to distractors, and maintain correctness across connected tool-use trajectories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.