Multi-Agent Debate Under Shared Bias: Beyond Re-Elicitation
Abstract
When language-model agents share a biasing context they can fail together, and their agreement stops signaling reliability. Standard evaluations cannot say what multi-agent debate adds in this regime: debate-vs-single-pass comparisons confound argument exchange with fresh re-elicitation, and clean-verifier comparisons confound the verifier's identity with its information position. We separate both confounds by matched intervention on induced false-consensus states, where at most two of five initial committee votes are correct. An equal-budget control re-elicits every agent afresh each round and withholds only peer arguments; a same-verifier intervention asks one held-out verifier the same questions with and without the biasing context. We run the decomposition under two mechanistically different shared-bias inductions: an authoritative premise, and then a preregistered replication with corrupted retrieved evidence that removes a residual asymmetry of the first. Peer argument exchange improves recovery by pp (95% CI []) beyond equal-budget re-elicitation, and context exposure alone costs the same verifier pp on average, with strong task heterogeneity; both preregistered margins clear. Exchange recovers states in which no initial vote was correct, so the effect is not amplification of a correct minority. What does not carry over is debate's ranking against a context-matched single verifier, which depends on the induction. In a simple shared-bias/private-evidence model we characterize when unanimously wrong votes conceal collectively sufficient evidence and when one-bit vote compression strictly loses Bayes accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.