Hard Enough to Learn: Evolving Verified Tool-Use Environments with Code Agents
Abstract
Large language model (LLM) agents have emerged as a dominant paradigm in artificial intelligence, with their capabilities increasingly determined by their ability to invoke and orchestrate external tools. Reinforcement learning acquires that ability at scale, and it needs verified environments with evolving difficulty to keep the training signal rich. We build a factory that produces verified RLVR data with code agents. An agent authors a domain, tools with implementations, a seeded database, and declarative task templates. Deterministic code then instantiates each task, derives its verified gold tool-call chain. One deterministic predicate validates the gold, measures each task's difficulty, and serves as the reward, so no judge sits in the loop. The factory remains effective across different code agents and harnesses. Per-task difficulty estimates calibrate the corpus to a learnable band and let the code agents harden the saturated tasks. Reinforcement learning on the band-calibrated corpus improves Qwen3.5 tool use across scales. We evaluate these capabilities on GAIA2, WorkBench, -bench, BFCL, Workplace Assistant, and ToolSandbox, and assess capability retention on general benchmarks and IFBench. Training on this corpus outperforms training on randomly sampled task sets of equal and twice its size. The same corpus supports supervision beyond rewards. Supervised fine-tuning develops tool-use capabilities in weak non-Qwen base models from the Llama-3.2 and SmolLM families. We also demonstrate knowledge transfer through on-policy distillation within the generated environments. The 4B student achieves approximately 98% of the average performance of its 9B teacher across the evaluated benchmarks, while the 2B student outperforms a baseline trained with supervised fine-tuning followed by GRPO.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.