StateForge: Synthesizing Verifiable Long-Horizon Data for Tool-Using Agents via State-Change Supervision
Abstract
Synthesizing long-horizon training data for tool-using agents requires reachable goals and verifiable outcomes. Successful tool calls do not guarantee task completion, while reference-call matching can reject valid alternative executions. We propose StateForge, a framework that unifies data synthesis, supervised fine-tuning (SFT), and reinforcement learning (RL) through state-change supervision. StateForge constructs executable environments with explicit state spaces and tools for querying and transforming states. Executing dependency-guided tool workflows yields reachable state-change targets, which ground task synthesis and task–target consistency checks. For each task, three trajectory synthesis modes preserve the same target: Direct provides complete goals, Progressive reveals goals in stages, and Recovery introduces recoverable environment perturbations. This shared target supports exact state-change matching to filter SFT trajectories and RL rewards that credit required changes and penalize unintended modifications. StateForge produces 544 environments and 1,697 tasks, with an average of 69 tool calls per task. With only 4B parameters, StateForge achieves state-of-the-art performance among specialized agent models on representative tool-use benchmarks, attaining 91.17% macro-average success on -Bench, 43.33% accuracy on DeepPlanning Shopping, and a 27.25% overall score on BFCL-v3 Multi-Turn.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.