acceptodds
Under review as a conference paper at ICLR 2027

Solver-Guided Reasoning for Mixed-Equilibrium Strategies

Abstract

Reasoning in large language models (LLMs) is often grounded in human text, human demonstrations, and human-generated rationales. For equilibrium reasoning in complex games, however, relying on human data can be suboptimal. Human play is often guided by intuition and heuristics, while human explanations emphasize individual actions rather than mixed strategies. To strengthen LLM reasoning in games, we study how to articulate equilibrium strategies using solver output. We propose a Mixed-Strategy Decision Tree (MDT), which represents solver mixtures through locally sparse, inspectable decision rules. We evaluate this representation through fixed-LLM strategy prediction, using Scenario-Constrained Counterfactual Sampling (SCCS) to construct same-context hand comparisons. We instantiate the study on No-Limit Texas Hold'em using over 250 million solver-labeled decisions. With identical reference policies, MDT routing reduces interpolation error by 9.8% relative to weighting by input distance on held-out boards. In-context comparisons across eight LLM configurations show improved prediction from solver-backed references. With forward-verified inputs and all requests counted, error falls by 37.7% for DeepSeek and 41.6% for GPT on policy-filtered targets; independently fixed, training-held-out queries yield reductions of 26.2% and 20.8%. Complete River policies and Liar's Dice extend the evaluation to strategic compression and a second game. The results establish useful policy relations in MDT routing and a practical interface for studying how LLMs use solver-backed evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.