Failure Migration in GUI Agents: Outcome Churn and the Trade-offs of Reflective Recovery
Abstract
Higher aggregate success can conceal task-level regressions and successful states lost during retry. We study these two forms of outcome change and their resource trade-offs through paired OSWorld evaluations, same-configuration reruns, and reflective recovery. Scaling Qwen3.5 from 4B to 9B raises recorded full-split success from 21.4% to 26.3%. Among 342 tasks with valid evaluations for both models, 35 improve and 22 regress; an always-on accessibility scaffold yields 21 improvements and 31 regressions on 339 evaluable pairs. Same-configuration reruns also exhibit substantial outcome churn. We evaluate grounded verification and reflective retry with an always-strong agent and FMART, a self-hosted 4B agent with failure-triggered strong-tier takeover. The always-strong reflective system improves full-split success from 51.5% to 59.1% on 369 tasks; its gain remains significant in a runner-matched paired comparison. With a three-round cap, FMART reaches 57.2%; its retained records contain 0.77 times as many strong execution steps as the always-strong reflective system. Both record approximately USD 0.42 per included task in their cost ledgers, which combine API estimates with nominal local-step charges for the cascade. In two FMART runs, five and seven tasks, respectively, have successful checkpoints followed by zero-score selected endpoints. Each of these tasks contains a successful checkpoint labeled incomplete by the completion checker. Together, these results show why evaluating GUI-agent improvement requires paired outcomes, returned-state success, and separate execution and monetary resource measures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.