CUWorldEval: Learning Worlds for Multi-Turn Computer-Use Agent Evaluation
Abstract
Evaluating computer-use agents requires knowing whether they can complete tasks through interaction, not just match recorded actions. Offline action matching is inexpensive but cannot capture the consequences of alternative actions, while evaluation in real applications is costly to repeat. We introduce CUWorldEval, a learned environment that predicts screen descriptions and available controls after each action, allowing agents to attempt multi-turn tasks without executing the underlying applications. The key challenge is that plausible individual predictions do not guarantee reliable task evaluation. Errors can accumulate, change the agent's subsequent actions, and lead to task outcomes that differ from real software. We propose Multi-Turn On-Policy Distillation (Multi-Turn OPD), which combines error-conditioned teacher training with distillation on student-generated interaction histories. We first train a teacher to predict reference next states from the student's imperfect histories, using full reference-trajectory summaries to guide correction. We then freeze the teacher and distill it on fresh student rollouts under recorded actions. Each predicted state becomes part of the next turn's history, training the student to continue despite its own errors rather than only from accurate reference states. We evaluate the same agents on matched Word, Excel, and PowerPoint tasks in CUWorldEval and real software. For the two agents with offline baselines, CUWorldEval yields higher balanced accuracy, and it reaches up to 76% agreement with real-software task outcomes. These results support learned environments for interactive agent evaluation and show why their reliability must be assessed through complete task outcomes, not only individual predictions. Code is available at https://anonymous.4open.science/r/CUWorldEval-05EB.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.