RIVER: Recursive Self-Improvement of LLM Poker Agents
Abstract
Agent skills, natural-language instructions written by human experts, are increasingly used to extend large language model (LLM) agents without retraining, and have recently been applied to strategic games such as poker. For a recent expert-written skill for heads-up no-limit Texas hold'em (PokerSkill), we find that the return varies considerably across base models, across reasoning efforts of a single model, and even across requests for more detailed analysis before acting that leave the encoded expertise unchanged. More reasoning thus does not guarantee better play, and the value of a skill depends on how each model executes it. Since we cannot know in advance how a model will apply fixed human expertise, this raises the question of whether a model can start from human expertise and iteratively derive a poker skill suited to itself. Poker makes such self-improvement difficult, since single-hand outcomes are highly variable with no verifiable reward for individual decisions, and the decision space is too large for hand-written advice to cover. We introduce (\riverexpansion), in which the LLM that plays poker recursively improves its own skill library. Each update reflects retrospectively on three complementary evidence streams, namely the human expertise encoded in the original skill, the model's execution of its current skill, and a comparative analysis of returns across skill versions. We further factorize the skill by decision factor and aggregate evidence from single decisions to similar states, condensing the reasons behind them into updates of the skill. Across four configurations spanning three model families, derives for every model a skill that reaches a positive mean return against the Slumbot poker bot and exceeds the expert-written skill by 383 to 526 milli-big-blinds per hand.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.