RSI-CUA: Towards Recursive Self-Improvement of Computer-Use Agents in an Endless Training Ground
Abstract
Computer-use agents (CUAs) can improve through environment interaction, but simply accumulating more experience does not determine which experiences best drive progress. As an agent's mastered capabilities and their compositions continually evolve, the central question for sustained improvement is: what should the agent learn next? We present RSI-CUA, a framework that drives recursive self-improvement through an explicit, evolving capability state. In each stage, trajectory-grounded capability profiling analyzes task semantics, interaction traces, and execution outcomes to distinguish capabilities that were successfully exercised, attempted but failed, or unreached. Guided by these identified gaps, RSI-CUA synthesizes new executable tasks calibrated to the learner's empirical capability boundary by balancing complexity, horizon, and composition. Validated tasks accumulate into an expanding Endless Training Ground for policy optimization via rejection-sampling SFT and online RL with GRPO, while a fixed 400-task profiling benchmark, GeneralCUABench, tracks stage-consistent capability progression. Evaluated on Qwen3-VL-8B-Instruct across three successive stages, pass@1 success on GeneralCUABench rises from 3.25% to 8.50%, 11.25%, and 17.75%, with the capability-guided curriculum outperforming an unguided counterpart under a matched training budget. On OSWorld, held out from profiling, synthesis, and policy optimization, the score improves from 18.06% to 25.28%. Capability analysis confirms that these gains stem from better action composition, robust state tracking, artifact persistence, and reliable task termination.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.