acceptodds
Under review as a conference paper at ICLR 2027

Coach-Guided Autonomous Reflection Improves Deception and Detection in Multi‑Agent Werewolf Games

Abstract

Previous studies have demonstrated that large language models (LLMs) are capable of generating deceptive statements in interactive settings, whereas their ability to detect deception and improve these capabilities through experience remains comparatively under-explored. We present an autonomous reflection framework for multi-agent Werewolf games that enables LLM agents to extract, organize, and reuse strategic skills for both deception generation and detection. The framework introduces a Coach Agent that automatically extracts reusable skills from gameplay trajectories, removes redundant or low-quality strategies, and dynamically injects curated skills into subsequent games. To assess strategic competence beyond game outcomes, we further propose a dual-path game intelligence evaluation framework that combines Monte Carlo simulation with LLM-based judgment. Empirical evaluations across baseline games demonstrate a balanced win rate with 78.3% deception detection accuracy. Coach-guided skill learning improves the win rate of the enhanced faction to 65–66%, while maintaining balanced competition when both factions are enhanced and yielding consistently higher game intelligence scores. Cross-model experiments reveal that while detection enhancement transfers robustly across various model architectures, deception production skills show model-dependent performance. Overall, the proposed framework complements win-rate-based evaluation and establishes a practical paradigm for continual strategic learning and game intelligence evaluation in complex multi-agent environments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.