Whose Outcome Is the Team's Score? Auditing Aggregation and Scored-Object Binding in Multi-Agent Evaluation
Abstract
Multi-agent systems increasingly coordinate multiple agents to solve complex tasks, yet their evaluation often compresses several agent executions into a single team score. Interpreting such a score requires two mappings that are rarely made explicit: how agent-level outcomes are aggregated into a team outcome, and which execution state or artifact is actually consumed by the grader. We formalize these as agent-to-team aggregation and scored-object binding, and introduce an executable aggregation receipt that makes this measurement path auditable from execution evidence. We apply the framework to a released multi-agent evaluation harness using source-code audits, controlled interventions, and model-driven runs. We then test whether the same audit transfers to an independently developed multi-agent framework. We find that nominally separate agents can become coupled through shared mutable state and that a reported team score can depend entirely on a single agent. In controlled interventions, changing the score-bound agent changes the reported result in all 24 paired comparisons, whereas changing either of the other agents changes none. We further show that unspecified aggregation creates a deterministic ambiguity whose width equals the agent-disagreement rate. On the independent framework, a source-pinned native adapter reconstructs the executed decision path while revealing a mismatch between the declared consensus semantics and the recorded accounting. Finally, we repair the harness by isolating agent states, explicitly aggregating all agent outcomes, and preserving scored-object lineage. Our results show that reliable multi-agent evaluation requires an auditable path from individual agent outcomes to the object actually scored.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.