acceptodds
Under review as a conference paper at ICLR 2027

Terminal-World: Scaling Terminal Agent Environments via Agent Skills

Abstract

Terminal agents enable Large Language Models (LLMs) to flexibly compose command line actions, extending their action space to achieve superior generalizability. To train such agents, existing methods construct training data by first fixing one of the three required components (i.e., a task instruction, an executable environment, and an expert trajectory), such as tasks from seed keywords or environments from GitHub repositories, and then completing the rest around it. However, such pipelines synthesize the three components in isolation, making it difficult to maintain coherence among the three as data scales, limiting the effectiveness of data scaling for agent improvement. To mitigate this, we introduce Terminal-World, a unified framework that uses agent skills as its core primitive to automatically synthesize coherent training data at scale. Specifically, each skill naturally encodes what to accomplish, when to apply it, and how to execute it, enabling the task, environment, and trajectory to be jointly derived from a single specification rather than assembled independently. With this framework, we construct the dataset Terminal-Atlas and train Terminal-Navigator-8B/14B/32B. Using the same teacher and base model with only 1.2% of the training data, Terminal-Navigator-32B achieves 31.5 vs. 27.0 Pass@1 and 43.8 vs. 37.1 Pass@3 on Terminal-Bench 2.0. These results highlight Terminal-World's ability to efficiently synthesize high quality training data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.