acceptodds
Under review as a conference paper at ICLR 2027

Tmax: A simple recipe for terminal agents

Abstract

Terminal-using agents have quickly become a popular and important downstream application of language models (LMs). Despite their prevalence, relatively little academic work has examined RL-based training of these models, due to difficult benchmarks, a lack of data, and a lack of simple baseline recipes. We present TMAX, an open RL recipe for terminal agents that brings small open models closer to the frontier. While simple, our recipe achieves 33% on Terminal-Bench 2.0 with only 9B parameters, outperforming much larger models from prior work. Our recipe first involves generating data using a novel taxonomy, combining difficulty control, personas, and verifier diversification, which allows us to cheaply generate large amounts of terminal environments for RL and SFT training. We then train open weight models using RL with our data, using a simple, outcome-only recipe, with small adjustments to improve stability. Our recipe works for models from 2B to 27B and across model families, dominating the Pareto curve for open terminal-agent performance at smaller sizes. We will publicly release our data, models, and code as a strong baseline for future open academic work on terminal agents

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.