Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate
Abstract
Multi-agent debate (MAD) can improve individual answers, but convergence does not by itself show that peer information helped. We separate answer changes into reconsideration instability under information isolation, strict conformity, and reasoning-only peer adoption. In the primary MMLU-Pro setting, 37% of agent-question pairs change under self-reflection alone. Among strict-conformity events, 30.7% are correct-to-wrong, 36.4% are wrong-to-correct, and 32.9% remain wrong. These directions vary with correct-answer availability and peer support. Individual accuracy rises from 36.9% at Round 0 to 45.9% under full reasoning, while majority-vote accuracy is unchanged between self-reflection and full reasoning (77.35%). A controlled, susceptibility-selected information-gradient experiment with an assigned correct prior finds that invalid and question-specific wrong reasoning increase wrong-answer adoption, with the invalid-reasoning contrast reversing across target models. A diagnostic model predicts harmful conformity only in an oracle-eligible setting (AUC = 0.784). An answer-agnostic model predicts peer adoption (AUC = 0.804), but the targeted intervention did not produce a detectable accuracy gain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.