When Does Multi-Agent Debate Backfire? A Predictable and Preventable Failure Mode of Heterogeneous LLM Debate
Abstract
Multi-agent debate (MAD) is widely believed to improve the factuality and reasoning of large language models (LLMs). We show that in heterogeneous debate, the realistic setting where a strong agent debates weaker ones, debate can systematically destroy correct answers. Across three strong models (Claude-Opus, GPT-5.2, and Gemini-3; four model configurations over 4,554 calibrated BBH tasks), we find that on questions the strong agent initially answers correctly, debate flips 11.6% of them to wrong (a harmful-debate event), with rates as high as 23.6%. Crucially, this failure is not random but predictable before any debate round is spent: whether the strong and weak agents agree in round 0 almost perfectly separates safe from harmful debate. The harmful rate is 0.2% under round-0 consensus versus 26.4% under round-0 disagreement (a 129-fold gap, p < 10^-125 when merged; significant for every strong model individually). A single zero-cost observable feature, round-0 disagreement, recovers harmful debate with approximately 100% recall. We further show the failure is preventable: a simple training-free rule, PROTECT, that defers to the strong agent's initial answer whenever agents disagree, recovers accuracy on the at-risk subset to approximately 100% while skipping approximately 50% of debate compute; several variants even exceed the strong-solo baseline. We release the benchmark, predictability protocol, and PROTECT baselines as a testbed for when to trust debate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.