Rewind and Retry: On the Test-time Scaling and Harness of Long Horizonal Computer Use Agent
Abstract
Computer-use agents (CUAs) interact with computers through GUIs and CLIs to complete complex tasks. As task horizons grow longer, the bottleneck shifts from the policy itself to the execution harness and test-time scaling strategy. We study this harness axis rather than proposing a new policy. Simple test-time scaling methods, such as best-of-N, resemble depth-first search and are ineffi- cient in exploring the solution space. We study three mechanisms for test-time scaling in computer-use agents: pruned tree search, evidence-based evalua- tion, and hybrid GUI–CLI execution. We instantiate these ideas in a practical harness, Checkpoint Tree Search (CTS), which decomposes a task into mile- stones, branches from VM checkpoints, selects among candidate continuations using evidence-based verification, and resumes execution from the selected state. Against a single rollout in the same harness, CTS improves mean graded score by +9% relative on OSWorld and +48% relative on the longer OSWorld-V2. On the multi-milestone subset of OSWorld-V2, three-branch CTS exceeds the judge-free pass@3 upper bound of independent full-episode sampling best-of-3, suggesting that enabling mid-episode rewinds breaches the theoretical ceiling of independent trajectory sampling. This work also presents extensive and novel empirical studies on the harness and evaluation for CUAs. We will release the complete reproduce- able codebase on GitHub.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.