CoordLedger: The Coordination Tax of Multi-Agent Systems
Abstract
Multi-agent LLM systems are usually judged by whether they outperform a single agent. Yet a multi-agent win can come from three different sources—more inference calls, access to more information, or better coordination—and standard comparisons cannot tell them apart. We introduce , a benchmark that compares each communicating system against the single-agent references implied by its deployment access contract. When all task-critical information could be given to one model, the deployment is centralizable, and the reference is a compute-matched single model. When no single role may see all of it, the deployment is structurally distributed, and we use two references: a single agent under the same access restriction, and a compute-matched single agent that sees everything. contains 1,957 instances across five tracks, with 58,248 frozen core runs over six model configurations from five model families. The two regimes behave very differently. On centralizable tasks, the compute-matched single model matches or beats both the Star and Chain topologies in five of six configurations. On structurally distributed tasks, communication clearly helps: Chain raises mean accuracy from 0.188 for the access-restricted agent to 0.335. But the all-access agent reaches 0.495, so communication recovers less than half (48%) of the value that pooling information makes available. We call the remaining gap the coordination tax. Stage-level traces show where it is paid: as information becomes distributed, the dominant failure shifts from delivering evidence to integrating it at the final decision. A multi-agent gain, in short, is only interpretable against the reference its deployment permits.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.