acceptodds
Under review as a conference paper at ICLR 2027

When Does Multi-Agent Debate Backfire? A Predictable and Preventable Failure Mode of Heterogeneous LLM Debate

Abstract

Multi-agent debate (MAD) is widely believed to improve the factuality and reasoning of large language models (LLMs). We show that in heterogeneous debate, the realistic setting where a strong agent debates weaker ones, debate can systematically destroy correct answers. Across three strong models (Claude-Opus, GPT-5.2, and Gemini-3; four model configurations over 4,554 calibrated BBH tasks), we find that on questions the strong agent initially answers correctly, debate flips 11.6% of them to wrong (a harmful-debate event), with rates as high as 23.6%. Crucially, this failure is not random but predictable before any debate round is spent: whether the strong and weak agents agree in round 0 almost perfectly separates safe from harmful debate. The harmful rate is 0.2% under round-0 consensus versus 26.4% under round-0 disagreement (a 129-fold gap, p < 10^-125 when merged; significant for every strong model individually). A single zero-cost observable feature, round-0 disagreement, recovers harmful debate with approximately 100% recall. We further show the failure is preventable: a simple training-free rule, PROTECT, that defers to the strong agent's initial answer whenever agents disagree, recovers accuracy on the at-risk subset to approximately 100% while skipping approximately 50% of debate compute; several variants even exceed the strong-solo baseline. We release the benchmark, predictability protocol, and PROTECT baselines as a testbed for when to trust debate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.