TraceGym: Failure-Driven Environment Synthesis for Training Terminal Agents
Abstract
LLM agents are increasingly used to solve tasks in terminal environments, yet they still make frequent mistakes. These failures are a valuable learning signal, because the same kinds of errors, such as declaring success without checking the output, recur across otherwise unrelated tasks. Prior work uses such failures as feedback at inference time, fine-tunes on corrected trajectories, or synthesizes new tasks around the capability that failed. However, none of them makes the agent revisit the situation in which it erred and learn to act correctly there. We introduce **TraceGym**, a fully automated pipeline that turns an agent's failures into executable training environments, each placing the agent back in a situation where it made a wrong decision so that it can learn the right one. It has three stages. *Failure diagnosis* distills the agent's errors into *failure contracts*, reusable specifications of a wrong decision and its correct alternative. *Environment synthesis* recreates each contract in tasks drawn from a separate training pool, through a minimal change that leaves the task's objective and verifier intact. *Training* then uses this existing verifier as the only reward, so no new reward design is needed. From failures of Qwen3.6-27B on Terminal-Bench 2.1 and a task pool disjoint from the benchmark, TraceGym produces 225 training environments. Training the same model on these environments, without a stronger teacher, raises its Terminal-Bench 2.1 success from 48.8% to 56.2%, compared with a 4.4-point gain reported for reinforcement learning on the same pool. The gain generalizes to TerminalWorld-Verified, an out-of-distribution benchmark, where success rises from 32.0% to 37.5%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.