EvoTerminal: Co-Evolving Terminal Agents through Scalable Environment Evolution
Abstract
Sustained self-improvement requires not only updating the learner, but keeping its training experiences useful as capabilities change. Executable terminal environments provide a scalable and verifiable substrate for such practice, yet scaling their supply alone cannot keep the training distribution aligned with an evolving learner: informative tasks become redundant, previously unreachable requirements become learnable, and new capability bottlenecks emerge. Crucially, this mismatch is not reducible to difficulty, since a task can remain appropriately challenging while exercising behavior that is no longer limiting. We therefore argue that the executable training distribution should evolve with the learner, adapting both what to practice and how hard. We introduce EvoTerminal, a self-improvement framework that co-evolves terminal agents with their executable training distributions. Its Discover–Shape–Evolve loop identifies the learner's behavioral frontier from feedback, reallocates practice across capabilities while calibrating difficulty, and realizes these targets as diverse environments that preserve the learning objective. The updated learner reshapes the next distribution, making environment evolution a recurring part of training rather than a one-shot data-generation stage. The same loop supports both model and harness adaptation. On Terminal-Bench 2.0, EvoTerminal raises Qwen3-8B SFT by 12.0 points and Qwen3.6-35B-A3B skill optimization by 5.2 points over their bases, with gains transferring across benchmarks and backbones. Ablations further confirm the importance of feedback refresh and joint capability–difficulty adaptation, underscoring the value of co-evolving training distributions with the learner.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.