MAScope: Diagnosing Collaboration Loss in Multi-Agent Systems
Abstract
Multi-agent systems assemble agents with different expertise, yet performance depends on how that expertise is organized, not on how much of it is present. Existing benchmarks only evaluate a method's outcome or a prescribed team's process, and cannot tell what a method's own collaboration contributes. We introduce **MAScope**, an environment of 51 expert agents holding private knowledge from Stack Exchange and 2,766 collaborative tasks with verifiable evidence and work dependencies. MAScope follows each required contribution from the agent that produces it into the final answer, splitting collaboration into four measurable capabilities: *expert selection*, *task orchestration*, *information transfer*, and *result integration*. Comprehensive experiments on twelve representative methods show that having the required experts does not deliver the task. Once work turns interdependent, success drops by 48%, over 3× that of a no-collaboration single agent. Trajectory analysis attributes this loss to three gaps: (1) **planning rigidity** in task orchestration, (2) **semantic compression** in information transfer, and (3) **adjudicative aggregation** in result integration. Even exhaustive collaboration and added memory recover little of it, since neither changes how a method organizes the work. MAScope gives a common basis for measuring and improving multi-agent collaboration, available at https://anonymous.4open.science/r/MAScope-7B76.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.