Scaling Tool-Use Agents via Environment-Grounded Query Synthesis
Abstract
Training tool-use agents with agentic reinforcement learning (Agentic RL) is bottlenecked by two challenges: the lack of scalable, robust execution environments and the scarcity of realistic training data that carries learning signal and captures implicit human reasoning. Existing pipelines build the two separately from pre-collected tool corpora and overlook the difficulty and realism of synthesized tasks, limiting their effectiveness for RL. We introduce , a fully automated framework that addresses both challenges. autonomously explores and verifies stateful, executable tool environments from authentic resources, and synthesizes natural multi-turn trajectories through topology-aware sampling and implicitness injection, producing grounded queries with realistic intents. Using 85 environments across 7 domains, generates 2,575 SFT and RL trajectories and beats prior work with significantly fewer environments and data. Across Qwen3 models from 1.7B to 8B, improves over the base models by up to **15.0 points** on BFCL-v3 multi-turn, **8.6** on MCP-Atlas, **4.9** on -Bench, and **8.3** on VitaBench. By automating both environment construction and trajectory synthesis, offers a scalable, extensible foundation for Agentic RL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.