acceptodds
Under review as a conference paper at ICLR 2027

Who Broke the Debate? Training-Free Detection of Adversarial Agents via Token Confidence

Abstract

Multi-Agent Debate (MAD) systems improve reasoning performance through inter-agent collaboration, but they are also vulnerable to adversarially manipulated agents that can propagate misinformation throughout the debate process. Existing defenses rely on task-specific training, external LLM judges, or post-hoc communication analysis. These approaches limit cross-task generalization, increase computational cost, and delay intervention until adversarial information has already entered the debate. In this paper, we propose DICE, a training-free defense method for proactively detecting compromised agents in MAD systems using LLM internal signals. Our key insight is that adversarial prompts can induce systematic shifts in token-level confidence patterns during generation, particularly in the early stage. Based on this observation, DICE measures inter-agent confidence differences using token-level log-probability signals without requiring additional training or external supervision. Our experimental results show that DICE achieves strong and consistent defense performance across diverse MAD tasks. The data and code used in this work are available at: https://anonymous.4open.science/r/TtJd5Tu2zX48ISYkkWsd

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.