Coach-Guided Autonomous Reflection Improves Deception and Detection in Multi‑Agent Werewolf Games
Abstract
Previous studies have demonstrated that large language models (LLMs) are capable of generating deceptive statements in interactive settings, whereas their ability to detect deception and improve these capabilities through experience remains comparatively under-explored. We present an autonomous reflection framework for multi-agent Werewolf games that enables LLM agents to extract, organize, and reuse strategic skills for both deception generation and detection. The framework introduces a Coach Agent that automatically extracts reusable skills from gameplay trajectories, removes redundant or low-quality strategies, and dynamically injects curated skills into subsequent games. To assess strategic competence beyond game outcomes, we further propose a dual-path game intelligence evaluation framework that combines Monte Carlo simulation with LLM-based judgment. Empirical evaluations across baseline games demonstrate a balanced win rate with 78.3% deception detection accuracy. Coach-guided skill learning improves the win rate of the enhanced faction to 65–66%, while maintaining balanced competition when both factions are enhanced and yielding consistently higher game intelligence scores. Cross-model experiments reveal that while detection enhancement transfers robustly across various model architectures, deception production skills show model-dependent performance. Overall, the proposed framework complements win-rate-based evaluation and establishes a practical paradigm for continual strategic learning and game intelligence evaluation in complex multi-agent environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.