Rethinking Data Generation for Long-Horizon Tool Calling
Abstract
Real-world tasks routinely require long sequences of tightly interdependent tool calls, yet existing benchmarks for evaluating agentic systems rarely reflect this complexity. While current benchmarks capture realistic multi-turn interactions and underspecified inputs, they focus on shallow execution trajectories with limited inter-function dependencies. We examine current data generation approaches for tool-calling benchmarks and argue that they inherently bias datasets toward low-complexity examples. We propose BRIT(Back translating and Refining Iteratively for Trajectory generation), which makes two systematic modifications compatible with existing pipelines: (1) generating trajectories before queries to optimize complexity explicitly, and (2) applying iterative refinement instead of filtering to preserve complex instances. We provide theoretical analysis showing two complementary results. First, filtering during post-processing inevitably produces a light-tailed complexity distribution, regardless of the complexity distribution of the trajectories initially generated. Second, when the generator already produces a heavy-tailed complexity distribution, per-node refinement budgets that grow logarithmically with trajectory complexity are sufficient to preserve that heavy tail after post-processing. Empirically, BRIT substantially increases the complexity of generated datapoints. Our work advocates for including high-complexity instances in tool-calling evaluation and provides practical, scalable methods for achieving this goal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.