acceptodds
Under review as a conference paper at ICLR 2027

Who Should Learn from Failure? PlayCode-RSI for Verified Co-Evolution of Code and GUI Agents

Abstract

Code agents can repair interactive software, but some defects only appear during execution. GUI agents provide runtime evidence, yet their failures may reflect their own perception, planning, or control mistakes rather than faulty code. This raises a central question: **who should learn from failure?** We study this question in interactive games, where editable code and repeatable interaction allow us to reproduce similar failures with different causes. Automatically verifiable outcomes then help determine whether the game, the player, or both need to change. We introduce **PlayCode-RSI**, a framework for verified co-evolution of code and GUI agents. Its core mechanism, CARE (Causal Attribution and Routed Evolution), combines visual trajectories, execution feedback, program states, and source code. It selects among five decisions: patch the game, update the player, update both, retry, or make no update. Candidate updates are replayed and must pass source-task, transfer, retention, safety, and cross-partner verification. Only verified code experience and GUI skills are inherited by later generations. We also introduce PlayAdaptBench, with Coding, GUI, and Joint RSI subsets, together with multi-turn vision–action–code training data. Evaluation separates partial progress from strict task completion and tracks both independent and joint capability growth. Experiments show sustained improvements in both agents and transfer to unseen tasks, with fewer incorrect updates and less partner-specific co-adaptation. These findings support failure attribution and verified inheritance as a basis for recursive self-improvement. The code, benchmark, and dataset will be publicly released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.