Who Broke the Debate? Training-Free Detection of Adversarial Agents via Token Confidence
Abstract
Multi-Agent Debate (MAD) systems improve reasoning performance through inter-agent collaboration, but they are also vulnerable to adversarially manipulated agents that can propagate misinformation throughout the debate process. Existing defenses rely on task-specific training, external LLM judges, or post-hoc communication analysis. These approaches limit cross-task generalization, increase computational cost, and delay intervention until adversarial information has already entered the debate. In this paper, we propose DICE, a training-free defense method for proactively detecting compromised agents in MAD systems using LLM internal signals. Our key insight is that adversarial prompts can induce systematic shifts in token-level confidence patterns during generation, particularly in the early stage. Based on this observation, DICE measures inter-agent confidence differences using token-level log-probability signals without requiring additional training or external supervision. Our experimental results show that DICE achieves strong and consistent defense performance across diverse MAD tasks. The data and code used in this work are available at: https://anonymous.4open.science/r/TtJd5Tu2zX48ISYkkWsd
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.