ScreenRSI: Recursive Self-Improvement for GUI Agent Harnesses
Abstract
Modern agents combine a large language model (LLM) with a harness that prompts the model and interacts with the environment to complete user tasks. Although this code layer strongly shapes agent behavior, improving it still relies largely on engineers inspecting failures and revising the system by hand. Harness-level recursive self-improvement (HRSI) automates this process by evaluating tasks, analyzing failure modes, and revising the harness over successive rounds while keeping the model and environment fixed. Its central challenge is that sparse, local failure evidence guides changes to shared logic that can affect many other tasks. Successive revisions may therefore accumulate useful behavior or repeatedly break existing capabilities. We study how local repairs can yield broader gains under a limited evaluation budget and whether these gains transfer to unseen tasks. We introduce ScreenRSI, an HRSI framework for GUI agents that uses a semantic task-similarity graph to select failure seeds for coverage and diversity and samples previously successful tasks as preservation tasks. A recurring cycle of task selection, harness repair, and controlled merging turns execution evidence into validated harness revisions. On MobileGym, with Qwen3.8-27B fixed as the acting model and optimization confined to training tasks, the GPT-5.6-Sol optimizer produces a harness with a peak test success rate of 74.77%, improving over the initial harness by 15.71% and reducing false completion by 10.55%. Under the same round budget, graph-guided selection improves test success by 5.73% and reduces false completion by 3.52% relative to random selection. Several key mechanisms for improving the harness emerged during the evolution of ScreenRSI, including state management, stuck detection and recovery, task-specific prompt guidance, and output-format validation. These findings provide insight guidance for designing stronger harnesses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.