Echoverse: Deep, Evolving Environments For Training Computer-Use Agents at Scale
Abstract
Training computer-use agents requires software application environments that the agents can interact with. Sandbox environments offer a better alternative to real ones, since they can be deterministic and safer to take actions in. However, creating these environments with sufficient depth and breadth of functionalities for training agents is non-trivial. We present Echoverse, a semi-automated pipeline that builds synthetic environments from specifications and target scenarios, with task outcomes verified against references derived from the environments' database. Echoverse works by co-evolving the environment and the model together in a loop. Specifically, for each step, we perform rollouts in our environments and use verifiers to discern between model successes/failures and environment issues. Based on the verifier analysis, we improve the environment and use the rollouts for training the model. By adjusting the environment with feedback from the rollouts, we are able to add functionalities needed to enable more tasks and increase the fidelity to real applications. We use Echoverse to produce a set of 10 environments covering different software use-cases. We show that fine-tuning on rollouts from environments verified only by unit tests and simple checks leads to poor transfer to real environments compared to fine-tuning on our environments (e.g., vs accuracy). % We show that fine-tuning on rollouts from "shallow” environments without further iterations leads to poor transfer to real environments compared to fine-tuning on "deep” environments from later iterations (e.g., vs accuracy). Building on these results, we introduce an RL training recipe that combines database grounded trajectory rewards with dense per-step feedback to address sparse supervision over long action sequences. We show that this recipe improves performance significantly in our synthetic environments (e.g., to pass@4).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.