GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models
Abstract
Fixed game suites for evaluating strategic reasoning in large language models (LLMs) saturate and are hard to keep free of contamination. We introduce GENSTRAT, a system for evaluating the strategic capabilities of LLMs driven by procedural generation of two-player zero-sum imperfect-information betting games, allowing an evaluator to draw novel, fresh games on demand. Each game is characterized on a scale of five complexity axes (state space, temporal contingency, opponent dependence, private information, precision) and measured by Monte Carlo simulation. Using this generator, we construct a 50-game benchmark and hold a tournament of twelve frontier and open-weight models, gathering data from over 40,000 LLM vs. LLM matches. We further construct a composite complexity axis, and find that, on the eight hardest games, the gap between the best and worst model is 1.7 times its value across the whole set of games (1.6 times with weights learned on held-out games). Ablations show that a higher thinking budget, disclosing the opponent's strategy, and higher reasoning matched to a game's specific demands (for private information and planning depth) were associated with higher margins for the LLM player.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.