acceptodds
Under review as a conference paper at ICLR 2027

RepThink: Self-Evolving Representations for LLM Minimax Reasoning in Adversarial Games

Abstract

LLMs possess extensive knowledge of games, yet often struggle to understand concrete states and reason about adversarial consequences. We address two coupled failures: failing to identify decision-relevant relations in game states and failing to compare actions by their worst-case consequences. A correct executable simulator alone is insufficient, because accurately simulated states may still be difficult for the LLM player to interpret for decision-making. We introduce RepThink, which self-evolves a representation skill that transforms game states into relational evidence and explains its meaning and applicability. The LLM player reuses this skill across current and simulated states within a MAX–MIN procedure, alternating between the two players’ objectives and comparing the worst examined continuations. After games, an evaluator agent diagnoses decision traces, revises the skill, and validates revisions by replaying the same positions. Across five board games, continued evolution improves performance under fixed-depth minimax reasoning. Against Stockfish limited to Elo 1320, RepThink with depth-two reasoning raises DeepSeek’s Chess win rate from 12% with Base ReAct to 42%. Our theory distinguishes evidence-interpretation errors from MAX–MIN reasoning errors and derives directional bounds on their effects on decision quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.