acceptodds
Under review as a conference paper at ICLR 2027

TerminalEvo: Evolving Environments at the Capability Frontier of Terminal Agents

Abstract

Synthetic data has become a central paradigm for scaling supervision in language model training. Extending this paradigm to terminal agents requires synthesizing executable environments where models can acquire experience through interaction. However, the learning value of generated environments is highly dependent on their alignment with the capabilities of the target agent: tasks beyond a model's current competence may consume substantial interaction budgets while providing limited learning signal. We introduce **TerminalEvo**, a framework for evolving terminal environments toward the empirical capability frontier of a target agent. Through iterative agent rollouts, TerminalEvo collects complementary feedback signals, where success rates estimate task difficulty and action-observation trajectories reveal behavioral limitations, missing dependencies, and opportunities for environment refinement. An LLM-based editor uses these signals to propose semantic modifications to task structures, constraints, and dependencies, which are then validated through executable contracts and automated verification before being incorporated into future training iterations. We investigate whether frontier-aligned environments enable more efficient learning under fixed interaction budgets and whether behavioral feedback improves environment evolution beyond outcome-only signals.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.