acceptodds
Under review as a conference paper at ICLR 2027

CUAEvolve: Self-Improving Computer-Use Agents for Evolving Applications and Workflows

Abstract

Computer-use agents (CUAs) can solve increasingly complex, long-horizon tasks, but application updates and evolving workflows can erode their acquired capabilities. Maintaining up-to-date training data requires frequent collection and curation, a costly and time-consuming process that is difficult to sustain. In this work, we introduce CUAEvolve, a computer-use agent that adapts its policy and reinforcement learning (RL) curriculum to application and workflow evolution through a self-improving loop. To improve robustness to visual and interface changes, CUAEvolve interleaves GUI actions with executable code that can consolidate multiple GUI operations into fewer coding sequences and bypass visual grounding when programmatic interfaces remain compatible across updates. To reduce environment data curation as applications and workflows evolve, we introduce a self-improving RL loop that uses failure diagnoses to synthesize targeted task instructions, initial environment states, and evaluators at scale, enabling continual policy improvement. Concretely, 1) application variants modify one interface property with surface appearance, geometry, naming, or interaction sequences; 2) workflow variants revise user instructions and introduce previously unseen workflows. To measure performance on evolved tasks, we introduce OSEvolveBench, which pairs original tasks with application and workflow changes. Experimental results show that the self-improving RL loop raises task scores on evolved applications and workflows to 25.0% and 22.3%, outperforming prior open-source CUAs at comparable scales. CUAEvolve also shows strong retention and cross-platform generalization, achieving 52.7% and 46.7% on OSWorld-Verified and WindowsAgentArena.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.