acceptodds
Under review as a conference paper at ICLR 2027

Multi-Agent Debate Amplifies but Rarely Creates: A Calibrated Dissection at the Frontier of Model Capability

Abstract

Multi-agent debate is widely deployed in the belief that interaction lets large language models (LLMs) reason beyond their individual capabilities, yet many benchmark–model combinations used to justify it are now saturated, obscuring how much of the reported gain comes from interaction itself. We rebuild the evidence base on a pre-registered difficulty window that retains only items on which agents genuinely disagree, and evaluate homogeneous multi-instance debate, a heterogeneous five-model panel, Asch-style confederate manipulations, and equal-budget replications of three role- and aggregation-based designs. Here debate converges quickly but improves slowly, and its consensus is often confidently wrong: agreement forms by the second round, yet up to 72.7% of a model's contested items end in a unanimously wrong consensus. What drives convergence depends on where an item sits in the panel's capability distribution: where someone can solve an item, correct answers propagate; where no one can, wrong answers are adopted 108 times against 4, and the strongest model becomes the largest source of adopted errors. Exchanging reasoning traces from the first round doubles adoption of the observed majority and halves minority survival, so the channel that repairs weaker members also accelerates consensus collapse on unsolvable items. Neither scripted saboteurs nor dedicated critics shifted these outcomes, and no replicated design improved significantly over plain heterogeneous debate. Debate therefore acts more as a selector over answers its members can already produce than as a creator of new ones. Progress may depend less on richer protocols than on panel coverage, staged exposure of reasoning, and knowing when consensus deserves trust.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.