acceptodds
Under review as a conference paper at ICLR 2027

From Intuition to Deliberation: Self-Evolving Reasoning for Strategic Card Games

Abstract

Strategic card games, characterized by partial observability and dynamic adversarial interactions, serve as critical benchmarks for evaluating the decision-making intelligence of Large Language Models (LLMs). Most LLM-based approaches prioritize policy performance while overlooking explicit reasoning generation. Achieving transparent decision-making typically relies on high-quality Chain-of-Thought (CoT) data; however, such annotations are scarce and costly in strategic card games, whereas prompt-generated rationales are often ungrounded or unreliable. Moreover, incorporating explicit reasoning into discrete decision-making tasks may incur generative overhead and risk performance degradation. To address these challenges, we propose Metis, a self-evolving framework designed to harmonize competitive proficiency with explicit strategic reasoning. Metis comprises three synergistic components: mixed Supervised Fine-Tuning (SFT), structured-reward Reinforcement Learning (RL), and an iterative self-distillation loop. Specifically, mixed SFT alongside structured-reward RL bootstraps initial policy-reasoning alignment and mitigates sparse supervision in discrete action spaces, while the iterative self-distillation loop anchors reasoning logic to high-confidence strategic actions, transforming autonomous expert-level decisions into reliable supervision signals. Extensive experiments on GuanDan and DouDiZhu demonstrate that Metis achieves a synergistic enhancement of competitive card-playing proficiency and explicit strategic reasoning. Our code is available at https://anonymous.4open.science/r/Metis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.