How Well Do LLMs Adapt Their Play? A Controlled Evaluation with Classic, Misère, and Novel Board Games
Abstract
Board games provide a controlled setting for testing whether large language models adapt their behavior when task objectives or rules change. We study two forms of adaptation within a General Game Playing (GGP) framework. First, classic-misère pairs invert the success criterion while preserving the board, actions, legality rules, and state transitions. Second, Centralis, an original game designed for this study, introduces a new spatial structure, action semantics, and victory condition. All actions are validated by a common GDL-based referee. We evaluate Qwen3-8B and Qwen3-14B in Instruct and Reasoning configurations against four calibrated GGP reference opponents. The Reasoning configurations show fewer model-generated forfeits and greater recovery from invalid responses, but differences in competitive performance depend on the game and opponent and require substantially more tokens and execution time. The classic-misère comparisons reveal an objective-adaptation gap: win rates decrease or remain low in several pairings after the success criterion is inverted, despite unchanged game mechanics. In contrast, Qwen3-14B Reasoning achieves high win rates against all four reference opponents in Centralis, while Qwen3-8B Reasoning generally completes matches but records few wins. These results distinguish adaptation to an inverted objective from play under a new rule system and show that operational robustness, competitive performance, and inference cost characterize different aspects of model behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.