RepThink: Self-Evolving Representations for LLM Minimax Reasoning in Adversarial Games
Abstract
LLMs possess extensive knowledge of games, yet often struggle to understand concrete states and reason about adversarial consequences. We address two coupled failures: failing to identify decision-relevant relations in game states and failing to compare actions by their worst-case consequences. A correct executable simulator alone is insufficient, because accurately simulated states may still be difficult for the LLM player to interpret for decision-making. We introduce RepThink, which self-evolves a representation skill that transforms game states into relational evidence and explains its meaning and applicability. The LLM player reuses this skill across current and simulated states within a MAX–MIN procedure, alternating between the two players’ objectives and comparing the worst examined continuations. After games, an evaluator agent diagnoses decision traces, revises the skill, and validates revisions by replaying the same positions. Across five board games, continued evolution improves performance under fixed-depth minimax reasoning. Against Stockfish limited to Elo 1320, RepThink with depth-two reasoning raises DeepSeek’s Chess win rate from 12% with Base ReAct to 42%. Our theory distinguishes evidence-interpretation errors from MAX–MIN reasoning errors and derives directional bounds on their effects on decision quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.