Finding Kaiji Among LLM Agents: When Everyone Yields
Abstract
When LLM agents compete for scarce resources, do they turn on each other, and if their societies fail, why? This study asks where betrayal should pay: restricted rock–paper–scissors, the elimination game of the manga Kaiji, with a closed deck. Twenty agents message each other freely; the rules never mention cooperation, and every readout comes from the engine log with no LLM judge. Agents negotiate tie pacts before 93.5% of matches and break only 10.2%, but the breaks are aimed: 70.0% play the one card that beats the agreed card (chance is 50%). Their limit is the division of roles: who proposes and who accepts. Removing memory removes proposers, and coordination falls to chance. Removing partner choice leaves agreements at 97.7% yet cuts coordination by two thirds: faced with conflicting proposals, both yield at once in 53% of cases and collide on the swapped cards, as two people do in a corridor. Under scarcity, eliminations rise and betrayals are aimed no better. Reasoning changes what the pact is for: with thinking enabled, the same model keeps the pact but breaks it three times as often, mostly with the beating card, and a single reasoning agent profits at its non-reasoning neighbours' expense. The convention is not universal. Seven open-weight models run without reasoning almost never form it: they play to win, fewer survive, and several raise the stakes; the one re-run with reasoning forms it in most matches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.