Does Multi-Agent Collaboration Pay When Agents Serve Users? Testing Agent Scaling in Dual-Control Settings
Abstract
Multi-agent systems (MAS) built from language models are increasingly deployed on the premise that collaboration among agents outperforms a single-agent system (SAS), yet when collaboration pays remains poorly understood. The most comprehensive study to date answers with a capability-saturation principle: coordination helps while the underlying SAS is weak and hurts once it is strong, with the crossing near 45% single-agent accuracy. That principle was derived on autonomous benchmarks, where no user takes part. We test it where agents increasingly operate, in dual-control settings, where the agent serves a user under a written policy and both act on a shared environment. On the airline, retail and telecom domains of -bench we compare a SAS with a centralized MAS built from the same model (an orchestrator, a database agent and a policy advisor) for six models from three families (four, from two families, in the primary analysis), over about 7,000 simulated conversations, with each system held to one completion-token ledger per episode. Under a fixed simulated user no cell shows a MAS benefit: the MAS is worse in 12 of 13 valid cells and tied in the thirteenth, a cell near 45% where the benchmark’s default simulator yields a 10-point MAS gain. Where the principle predicts a benefit, the MAS loses by 13 and 22 points, and removing the budget ceiling brings it at best level with the SAS at 5.5 times its median token spend. A capability index measured on held-out domains shows no detectable dependence of the gain on capability (slope , 95% CI ), although this interval is too wide to exclude a steep decline; the fitted gain is negative across the observed range, with an interval that excludes zero only from an index of 60% upward. The mechanism that would justify a MAS here is not detectable: the MAS shows no measurable reduction in wrongful database writes and converts other failures into budget exhaustion and inaction. For agents that serve people, these results make the SAS the default and place the burden of proof on multi-agent designs. We release all trajectories, budget ledgers, prompts and code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.