Generalizable Opponent Exploitation in LLM Poker Agents via Mixed Best-Response Training
Abstract
Opponent exploitation, adapting to and profiting from an opponent’s weaknesses, is a crucial capability in competitive games, yet existing LLM-based poker agents are mostly prompted or fine-tuned toward equilibrium play and fail to exploit suboptimal opponents. We introduce GOE-LLM (Generalizable Opponent Exploitation with LLMs), a framework for learning opponent exploitation in two-player zerosum poker. A lightweight MLP profiler summarizes an opponent’s recent behavior into a natural-language description, and an LLM exploiter conditioned on this description is trained with group relative policy optimization on a curated mixture of best-response strategies. We propose the Mixture-of-Best-Responses Principle, which excludes non-transitive dominance cycles from the mixture, to keep training stable while preventing collapse to a single equilibrium strategy. In Kuhn poker and Leduc Hold’em, GOE-LLM outperforms prompting and equilibrium-fine-tuned baselines across three model sizes and approaches the best-response upper bound, including against opponents unseen during training. Moreover, GOE-LLM trained only on Kuhn poker carries over zero-shot to unseen imperfect-information games in GTBench, improving the win rate by 18.0 points on average almost entirely through better strategy rather than fewer illegal moves, while bringing no gain in perfect-information games. These results suggest that mixed best-response training teaches LLM agents exploitation skills that generalize beyond the specific opponents and the single game seen in training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.