Verifiable State Transitions as Unified Supervision for Training Stateful Tool-Use Agents
Abstract
Training stateful tool-use agents requires evaluating not only whether they invoke the correct tools but also whether their actions produce the intended changes in persistent environments. However, supervision based primarily on final task outcomes does not explicitly assess intermediate execution, while tool-call matching alone cannot verify the actual effects of executed actions. Reliably assessing solution quality therefore requires jointly evaluating tool-call correctness and the resulting state transitions. We introduce VESTA (Verifiable Execution and State Transitions for Agents), a unified training framework that derives execution-grounded supervision from verifiable state transitions. VESTA constructs executable environments from API specifications and synthesizes tasks in an executable-first manner, recording reference tool calls and expected environment states at each task turn as state-transition contracts. These contracts enable joint verification of tool-call correctness and state changes, supporting SFT trajectory selection, policy-conditioned RL task selection, and RL reward computation without learned or LLM-based reward judges. Experiments across three Qwen3 model scales on BFCL-v3 Multi-Turn, ACEBench-Agent, and -Bench show improvements over the corresponding base models in all nine model–benchmark comparisons, with RL further improving over SFT in every setting. On Qwen3-4B, VESTA achieves absolute gains of 12.63, 20.00, and 8.76 points on the three benchmarks, respectively. Across model scales, VESTA outperforms EnvScaler in seven of nine comparisons and AWM in all six available comparisons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.