Scaling Capabilities and Competition in Strategic Interactions
Abstract
LLM agents are becoming autonomous and ubiquitous, moving us toward a world where agent-agent interactions carry real consequences for the people and institutions they represent. These interactions will be strategic, mixed-motive, and structurally diverse, involving agents that differ in capability and number. We introduce three new multi-turn negotiation games grounded in real-world scenarios, designed to enable experiment sweeps that systematically scale agent capability, number of agents, and degree of competition. Across runs and models, more capable agents reliably achieve higher payoff, with greater inequality in capability-diverse groups than homogeneous ones. For models with an Elo above , we observe the emergence of a ruthless coalitioning behaviour, where agents form the minimum supermajority needed to complete the negotiation while allocating zero payoff to excluded agents. Moreover, we find that a team of weaker agents can outcompete a strong agent, but the advantage degrades as the number of agents scales, as larger, more complex games magnify capability gaps and facilitate deceptive exploits such as falsely claiming team membership. Finally, we also scale test-time compute, finding negotiation payoff gains comparable to - Elo points in GPT-5 and Claude Sonnet 4.6.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.