Evaluating Agent Collaboration at Scale via Simulated Software Worlds
Abstract
Scaling agent collaboration requires understanding how agents choose work, coordinate changes, and manage shared resources. We introduce Software World, a framework for studying these decisions through the maintenance of real, interdependent Python packages. Software World has three components: an environment where agents maintain software together, a method for constructing worlds and evaluating changes on unseen packages, and infrastructure for running and recording simulations at scale. Agents choose their own tasks to improve software efficiency while preserving correctness. We use human-authored tests and benchmarks to measure correctness and speedups in dependent packages that agents never see during the simulation. Across 31 simulations in six-package worlds with four models, we collect nearly 14K agent work-session traces. Across all four model families, agents produce larger held-out speedups when they can collaborate than when they work in isolation. Claude Opus 5 agents frequently initiate exchanges, while GPT-5.6 Sol agents make no contact when communication is optional. Assigning each agent a package can improve performance compared with letting agents divide up the work themselves. When agents share a budget through a voting system, some fail to coordinate access, leaving shared resources unused after their individual budgets run out. We will release our framework and dataset upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.