Disentangling Agent Contribution in LLM Multi-Agent Systems through Controlled Interventions
Abstract
Identifying which agent actually drives the success of a large language model (LLM)-based multi-agent system remains an open challenge. Aggregate performance cannot answer this question, and our results show that attribution does not simply identify the most capable reasoner. We introduce MASContributionBench, covering seven architectures and seven task suites, and combine leave-one-out (LOO), sampled Shapley, and sampled Banzhaf attribution with controlled interventions on topology, role prompts, and runtime permissions. Complete MAS outperform single-agent and solo-role baselines, while the aggregate difference from random teams remains unresolved; collaboration gains alone therefore do not validate a prescribed organization. More strikingly, changing validated static predecessor sets reverses an unchanged position's mean LOO sign with modest team-utility change. Exchanging functional prompts moves credit to unedited roles without improving the team, demonstrating that contribution does not simply follow role function. Changing Coder write access leaves of outcomes unchanged, while granting final-answer authority to Verifier produces no Verifier-selected outputs in all runs. These findings reveal that measured contribution conflates useful reasoning, routing leverage, and control of the evaluator-visible output. Contribution therefore measures the counterfactual fragility of an execution protocol, rather than an intrinsic property of an agent or role.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.