acceptodds
Under review as a conference paper at ICLR 2027

OpenTerminal: Scalable Construction of Executable Tasks for Open Terminal Agents

Abstract

Terminal agents enable diverse real-world workflows, yet open-weight models remain unreliable on realistic terminal tasks. A key bottleneck is the scarcity of executable data. Each task must pair an instruction with a runnable environment, the fixtures it uses, and a verifier that evaluates whether the task is completed. Building such data at scale requires broad task coverage, cross-artifact consistency, and a way to recover from construction failures. OpenTerminal addresses these challenges through corpus-anchored synthesis, shared task contracts, and failure-directed repair. Task concepts are synthesized from 307,293 problem patterns across 36 public corpora, where a cluster-smoothed sampler flattens the sampling distribution over semantic regions and increases coverage of underrepresented areas. A hidden task contract resolves implementation choices left unspecified by the instruction and provides a shared specification for environment and verifier generation. When a validation failure persists, Diagnose–Route–Repair identifies a likely source stage, feeds diagnostic evidence, and regenerates the downstream artifacts, recovering tasks that would otherwise be discarded. The pipeline yields 11,713 executable tasks and 7,365 training trajectories. We fine-tune Qwen3.5-4B and Qwen3.5-9B on these trajectories, improving pass@1 over the base models by 56.8% and 63.1% on Terminal-Bench 2.1, and by 31.3% and 23.9% on OpenThoughts-TBLite. Extensive experiments and analyses confirm the contribution of each component and characterize the data the pipeline produces.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.