Terminal-World: Scaling Terminal Agent Environments via Agent Skills
Abstract
Terminal agents enable Large Language Models (LLMs) to flexibly compose command line actions, extending their action space to achieve superior generalizability. To train such agents, existing methods construct training data by first fixing one of the three required components (i.e., a task instruction, an executable environment, and an expert trajectory), such as tasks from seed keywords or environments from GitHub repositories, and then completing the rest around it. However, such pipelines synthesize the three components in isolation, making it difficult to maintain coherence among the three as data scales, limiting the effectiveness of data scaling for agent improvement. To mitigate this, we introduce Terminal-World, a unified framework that uses agent skills as its core primitive to automatically synthesize coherent training data at scale. Specifically, each skill naturally encodes what to accomplish, when to apply it, and how to execute it, enabling the task, environment, and trajectory to be jointly derived from a single specification rather than assembled independently. With this framework, we construct the dataset Terminal-Atlas and train Terminal-Navigator-8B/14B/32B. Using the same teacher and base model with only 1.2% of the training data, Terminal-Navigator-32B achieves 31.5 vs. 27.0 Pass@1 and 43.8 vs. 37.1 Pass@3 on Terminal-Bench 2.0. These results highlight Terminal-World's ability to efficiently synthesize high quality training data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.