Agents that Matter: How Subtle Choices Shape Agent Attribution in Multi-Agent Systems
Abstract
As multi-agent systems (MAS) become increasingly complex, identifying the contributions of agents is critical for system optimization. However, existing approaches lack a unified formulation for credit assignment. In this work, we formalize agent attribution as a cooperative game, parameterized by the coalition distribution, intervention protocol, and target metric. Using this framework, we demonstrate that intervention protocols induce distinct games: Agent ablation isolates structural bottlenecks, whereas introspective LLM judges fail to faithfully approximate this behavior. We further find that Leave-One-Out (LOO) identifies bottleneck agents as effectively as combinatorial methods at much lower computational cost. Treating prompt optimization as an intervention protocol, we show that prompt-conditioned attribution can identify optimization targets and improve MAS performance. Finally, we apply our framework to audit a medical MAS, revealing that agent contributions to diagnostic accuracy and ethical behavior are often decoupled. By intervening on counterproductive roles, we observe an increase in ethics alignment while maintaining diagnostic accuracy. Overall, this work provides a principled approach for cost-effective MAS attribution and intervention. Code for reproducing our experiments is available at https://anonymous.4open.science/r/agent-eval-B850
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.