Solver-Guided Reasoning for Mixed-Equilibrium Strategies
Abstract
Reasoning in large language models (LLMs) is often grounded in human text, human demonstrations, and human-generated rationales. For equilibrium reasoning in complex games, however, relying on human data can be suboptimal. Human play is often guided by intuition and heuristics, while human explanations emphasize individual actions rather than mixed strategies. To strengthen LLM reasoning in games, we study how to articulate equilibrium strategies using solver output. We propose a Mixed-Strategy Decision Tree (MDT), which represents solver mixtures through locally sparse, inspectable decision rules. We evaluate this representation through fixed-LLM strategy prediction, using Scenario-Constrained Counterfactual Sampling (SCCS) to construct same-context hand comparisons. We instantiate the study on No-Limit Texas Hold'em using over 250 million solver-labeled decisions. With identical reference policies, MDT routing reduces interpolation error by 9.8% relative to weighting by input distance on held-out boards. In-context comparisons across eight LLM configurations show improved prediction from solver-backed references. With forward-verified inputs and all requests counted, error falls by 37.7% for DeepSeek and 41.6% for GPT on policy-filtered targets; independently fixed, training-held-out queries yield reductions of 26.2% and 20.8%. Complete River policies and Liar's Dice extend the evaluation to strategic compression and a second game. The results establish useful policy relations in MDT routing and a practical interface for studying how LLMs use solver-backed evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.