AlphaExploitem: Going Beyong the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play
Abstract
Poker is an imperfect information game that has served as a long-standing benchmark for decision-making under uncertainty. To maximize utility beyond the Nash equilibrium, an agent can deviate from equilibrium policies to exploit suboptimal play. We introduce AlphaExploitem, which extends the well-performing RL poker agent AlphaHoldem with a hierarchical transformer encoder that reasons over previously played hands, and we train it to exploit a suboptimal play produced by learned playstyles. We use a style variational autoencoder that learns a continuous latent space of play styles from trajectory data and decodes any sampled latent into a queryable opponent policy. On Leduc, a standard benchmark for poker research, AlphaExploitem extracts big blinds per hand from the test-set, more than three times what a non-adaptive exact Nash equilibrium extracts from the same set, and roughly five times the league-only baseline. Because the generator learns its style space from general game histories, the same method can be applied on human hand histories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.