acceptodds
Under review as a conference paper at ICLR 2027

Probabilistic Consensus for Weighted Multi-Agent Debate

Abstract

Multi-agent debate (MAD) offers a promising approach to improving the reasoning capabilities of large language models (LLMs), but reaching a correct consensus requires accounting for uncertainty in both proposed answers and the feedback used to evaluate them. Agents may produce incorrect solutions, and their assessments of other agents’ solutions may themselves be mistaken. We introduce Weighted Iterative Society-of-Experts (WISE), a zero-shot framework for weighted multi-agent debate with probabilistic consensus. WISE augments textual critiques with discrete correctness weights assigned to candidate answers, providing structured evidence for modeling uncertainty in agents’ evaluations. At its core is WISE-Dawid–Skene, an unsupervised probabilistic model that jointly estimates solver and scorer error patterns and infers answer correctness from weighted responses across debate rounds. This formulation accounts for uncertainty in both answer generation and scoring, enabling consensus to reflect the accumulated evidence and inferred capabilities of participating agents. Iterative feedback guides solution refinement, while the joint probabilistic model integrates the resulting answers and assessments to infer a collective decision. Experiments on standard multimodal reasoning benchmarks demonstrate that WISE consistently improves zero-shot reasoning, outperforming leading models by 6–12% on average. These results highlight weighted debate and joint probabilistic modeling of answer and feedback uncertainty as an effective approach to improving collective reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.