SAGE: Seedless Agent Task Induction from Environments for Post-Training
Abstract
Training capable agents requires interaction trajectories grounded in executable environments, yet constructing such data for a new environment often requires manually designing tasks, organizing interactions, and verifying execution outcomes. We introduce SAGE, a framework for seedless construction of agent training data from environment specifications. SAGE generates requests from domain context, tool interfaces, user attributes, and initial states, without task-specific query seeds or reference solution paths. A persistent user generator and an agent engage in multi-turn interaction, while an isolated verifier checks candidate requests before execution and complete trajectories after execution. Rejected requests are removed and regenerated, and accepted trajectories are converted into supervised fine-tuning data. We use Qwen3.8-2.4T, Qwen3.8-Next-Flash, GLM5.2, and Kimi-K3 as teacher models to construct data for general tasks, tool interaction, and software engineering, and evaluate the resulting data against public-data baselines. Independent SFT runs on Qwen3.5-27B yield absolute score gains over the initial model of 1.11%, 9.25%, 0.86%, and 2.33% on General Avg, BFCL Multi-turn, Tau2Bench three-domain Overall, and SWE-bench Lite, respectively. The corresponding absolute gains over the best public-data baselines in our comparison are 10.82%, 19.75%, 7.50%, and 8.00%. These results show that SAGE can provide effective supervision for agent post-training across different environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.