acceptodds
Under review as a conference paper at ICLR 2027

Multi-Agent Computer Use

Abstract

Computer use agents (CUAs) today are primarily deployed as single serial agents. This setup is suboptimal for complex long-horizon tasks that benefit from task decomposition, parallel execution, and consistent re-planning based on new information. In this paper, we argue that we should instead move towards evaluating and building *multi-agent computer use* (MACU) systems. These systems, which emphasize planning and parallel execution, alleviate many of the shortcomings of single-agent CUAs. We propose a general multi-agent setup in which a manager model decomposes computer use tasks as a directed acyclic graph (DAG) of subtasks, encoding relevant dependencies and goals for subagents. At each iteration, the manager dispatches parallel CUA subagents to carry out nodes on the ready frontier of the DAG, and continuously revises the DAG (adding, canceling, or rewriting nodes) as new findings arrive from subagents. This design treats the partially observable environment of computer use as a first class challenge: information that downstream agents may not be able to re-observe are retained and passed forward through the manager and DAG structure. We demonstrate that MACU improves over strong single-agent baselines by on four computer use benchmarks, across desktop (OSWorld) and web navigation (Online-Mind2Web, WebTailBench, Odysseys) environments, with the largest and most significant gains on multi-step and long-horizon web tasks, and exhibits more favorable test-time scaling. On Odysseys, a long-horizon web navigation benchmark, MACU solves as many tasks as the single-agent baseline under the same per-task time limit (), and matches the final single-agent success rate sooner in end-to-end wall-clock time. Our findings highlight that multi-agent coordination is a promising axis for scaling computer use agents to work productively and more effectively for longer time horizons. We release all code and interactive visualizations at *removed_for_review*.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.