acceptodds
Under review as a conference paper at ICLR 2027

Agent Collectives Should Not Detect Their Own Imposters: A Chess Case Study

Abstract

A collective of AI agents collaborating on a task has the potential to outclass any individual agent for that task. We study the robustness of such collectives against possible imposters, i.e., agents that deliberately try to mislead their peers. Since a single imposter could undo the collective’s advantage, we need to detect them. We consider two strategies: (i) incorporate imposter detection into the participating agents, or (ii) consider a dedicated imposter detector outside the collective. We investigate this empirically on GAMBIT, a testbed in which 4 reasoning agents deliberate on chess moves. The setting is deliberately small, but still challenging for frontier models. We choose chess because a it allows objective and quantitative assessment (through a state-of-the-art chess engine) of both the gain of using a collective (vs. a single agent) and potential damage by imposters. Regarding the first strategy (i), we find that merely warning the agents of potential imposter presence is not benefical: it degrades decisions when no imposter is present, provokes reactions ranging from self-accusation to scapegoating, strongly inflates token use, and reveals to the imposter how its intentions were uncovered. We therefore recommend to follow strategy (ii) and use a detector that reads the collective’s deliberation but never joins it and only returns a verdict. Such a dedicated detector then faces the challenge of dealing with adaptive imposter strategies, which requires recalibration to new attack strategies after only seeing very few examples, rather than wait for a full retraining cycle. In our benchmark we show that a 3B language model with a meta-trained classification head (is agent i an imposter or not?) effectively achieves that: a single gradient step on 20 labeled examples suffices to adapt to an unseen imposter strategy. At matched zero-shot accuracy, this meta-trained detector yields 8× the adaptation gain of standard finetuning, at 14× lower training cost. To support further research on collectives that stay robust as their adversaries evolve, we release the GAMBIT benchmark, with 37,352 labeled deliberations spanning 240 evolved imposter strategies. Code and data are available at https://anonymous.4open.science/r/gambit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.