Probe-based Divergence Calibration for Safeguarding Heterogeneous Multi-agent Debate
Abstract
Topology-constrained multi-agent debate can aggregate fragmented evidence, but its safeguards face a fundamental ambiguity: under heterogeneous private evidence, benign disagreement can resemble adversarial deviation. Existing anomaly-based defenses may therefore isolate agents whose responses are unusual but legitimately evidence-driven. We introduce PDVC, which changes the coordinates in which disagreement is observed. PDVC constructs task-adaptive semantic probes from public task information and independent first-round responses, projects centered responses onto these probes to obtain discussion-relative conflict representations, and learns an agent-wise role readout for topology intervention. Our analysis shows that shared attack-associated variation induces structured low-rank conflict after projection, while post-hoc evidence controls reveal that maliciousness is characterized primarily by the direction rather than the magnitude of unexplained response variation. Across four multi-hop reasoning benchmarks, diverse attack structures, and multiple LLM backbones, PDVC improves malicious-agent detection while preserving substantially stronger downstream task performance and benign collaboration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.