acceptodds
Under review as a conference paper at ICLR 2027

Agents that Matter: How Subtle Choices Shape Agent Attribution in Multi-Agent Systems

Abstract

As multi-agent systems (MAS) become increasingly complex, identifying the contributions of agents is critical for system optimization. However, existing approaches lack a unified formulation for credit assignment. In this work, we formalize agent attribution as a cooperative game, parameterized by the coalition distribution, intervention protocol, and target metric. Using this framework, we demonstrate that intervention protocols induce distinct games: Agent ablation isolates structural bottlenecks, whereas introspective LLM judges fail to faithfully approximate this behavior. We further find that Leave-One-Out (LOO) identifies bottleneck agents as effectively as combinatorial methods at much lower computational cost. Treating prompt optimization as an intervention protocol, we show that prompt-conditioned attribution can identify optimization targets and improve MAS performance. Finally, we apply our framework to audit a medical MAS, revealing that agent contributions to diagnostic accuracy and ethical behavior are often decoupled. By intervening on counterproductive roles, we observe an increase in ethics alignment while maintaining diagnostic accuracy. Overall, this work provides a principled approach for cost-effective MAS attribution and intervention. Code for reproducing our experiments is available at https://anonymous.4open.science/r/agent-eval-B850

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.