acceptodds
Under review as a conference paper at ICLR 2027

Who Gets Blamed? Attribution Under Interference in Multi-Agent LLM Markets

Abstract

An audit of interacting LLM agents needs to distinguish the effect of changing an agent's own deployment from the effects of changing its peers. Market-level comparisons can combine these effects and give a different answer from an agent-level evaluation. We study this problem across eight open-weight models in a repeated Bertrand pricing market, using a profit-based collusion index as the outcome. A two-stage design varies the treated share and randomly selects recipients, separating own effects from spillovers. Estimated own effects have significantly different values between no-peer-exposure and all-peer-exposure settings. High self-play collusion scores also coexist with low win rates against other models. To assess a query-based audit, we obtain agents' reaction functions at fixed market states and compose the maps. The stochastic composition recovers the spillover direction in seven of eight tested model-identity cells, but performs poorly for prompt-level spillovers and own effects. These findings motivate deployment-specific safety evaluation: specify whose outcome is measured, report peer exposure, and check which effects a query-based audit can recover before relying on it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.