Fool Me Once: Scaling Interaction Reveals How LLMs Adapt to and Exploit One Another
Abstract
Large Language Models (LLMs) act strategically toward the models they interact with, adapting to an opponent's behavior and exploiting what it cannot see. We scale test-time interaction up to 100 turns in matrix games grounded in real-world harms, where LLMs play against fixed strategies and against each other. Before a history of interactions exists, models cooperate in 81.6% of first turns. After a single observed interaction, their cooperation rises to 92.3% against an always-cooperating opponent but falls to 45.8% against an always-defecting one. Furthermore, withholding a player's decision history from its opponent encourages exploitation in games that reward defection and increases cooperation in games where mutual cooperation creates the best outcome for both players. Over 100 turns of Chicken, withholding this history raises outcomes where the player defects while its opponent cooperates from 4.5% to 69.0% of turns, and telling the opponent that the history is hidden does not undo this. A model's first action therefore cannot reveal how it will respond to other models as interactions accumulate. Long-horizon behavior is already difficult to evaluate, and interacting models add strategic complexity that single-model evaluations miss. Thus, pre-deployment safety evaluations must test models against other models, including adversarial and competitive ones, over long interaction horizons.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.