acceptodds
Under review as a conference paper at ICLR 2027

No One Wins Forever: Rethinking Voting, Self-refinement, and Debate of LLM Agents

Abstract

Multi-agent debate (MAD) has been established as a paradigm for improving reasoning capabilities through communication among multiple large language models (LLMs). Despite its empirical appeal, the mechanisms underlying multi-agent debate remain underexplored. Existing work often draws largely binary conclusions, characterizing MAD as either effective or ineffective under narrow or even unfair experimental settings. Such analyses fail to capture the rich dynamics of debate, resulting in limited conclusions. In this paper, we develop a simple mathematical model that captures two properties of model behavior: generation quality and answer stability. Our theory characterizes the complete ordering among three methods: majority voting (MV), self-refinement (SR), and MAD. Instead of drawing a binary conclusion, our analysis suggests that when independent generation favors the correct answer, voting eventually dominates by amplifying this advantage. Otherwise, debate outperforms self-refinement when the correct answer is the most prevalent in the long-run group distribution. We further show that answer stability shapes this collective behavior: stronger interaction increasingly favors answers that the model is more likely to retain under revision. For a fixed number of agents, we also derive stationary accuracy limits and exact finite-budget evaluation formulas. Through extensive experiments across six benchmarks and six LLMs, we find that the proposed theory closely matches empirical results. Guided by these insights, we further propose gated debate, which decides whether to use debate or self-refinement at each turn based on a training-free statistic. Results show it outperforms both self-refine and debate baselines at larger model-call budgets. Code and data will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.