acceptodds
Under review as a conference paper at ICLR 2027

Assessing dynamic predictability in LLM strategic behavior

Abstract

Unpredictability is central to strategic behavior in competitive settings. Yet, evaluations of large language models (LLMs) in mixed-strategy games typically measure only the overall frequencies of actions, not whether those actions can be predicted from the history of play. Drawing on methods from behavioral game theory and decision neuroscience, we introduce the Protean Adversarial Choice Assay (PACA), a repeated matching-pennies paradigm in which an adaptive opponent detects and exploits statistical regularities in a model's action history. We evaluate seven LLMs from Anthropic, OpenAI, and Google under two framings of the same game: a minimal prompt that conceals the strategic context and elicits a choice framed as an outcome prediction, and a strategically informed prompt that explains the game and prescribes the equilibrium strategy of unpredictable 50/50 play. Informed prompting moves behavior farther from equilibrium, and win rates fall by 10.9 percentage points on average (45.6% down to 34.7%), with six of seven models showing a decline. Under this framing, the worst-performing models over-alternate most strongly despite maintaining near-balanced overall action frequencies. These findings show that assessment of equilibrium play should account for history-dependence, and that explicitly prescribing the equilibrium strategy can be counterproductive to its execution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.