acceptodds
Under review as a conference paper at ICLR 2027

Agent S4: Evolutionary Computer Use

Abstract

Computer-use agents are limited by the interfaces humans give them. An agent may type a spreadsheet cell by cell when a single line of code would do, or overwrite a file with a script while the application still holds unsaved edits. Both the interfaces and the knowledge of when to use each are built by hand, which cannot keep pace with new applications, workflows, and models. We introduce Agent S4, an autonomous computer-use scientist that studies computer-use agents and improves their interfaces using population-based evolution. A meta-agent examines task agents' trajectories on synthetic training tasks, without success labels, and writes new interfaces and usage knowledge into a repository. Population-based evolution then selects which repositories to develop further by their performance on held-out tasks the meta-agent never sees, favoring changes that generalize beyond the seen tasks. The evolved repositories contain interfaces no one specified, such as a single request that batches dozens of GUI actions, and observations that condense PDFs and archives into compact text. On OSWorld V2, Agent S4 with Qwen3.8-27B reaches 21.3% binary task success, on par with the 20.6% reported for Claude Opus 4.8 with batched actions, at an estimated task-solving cost of USD 1.10 per task versus the reported USD 72.40, approximately 1.5% of the cost. These results suggest that interfaces are an axis of improvement complementary to the model, one that can keep evolving as models and harnesses change.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.