acceptodds
Under review as a conference paper at ICLR 2027

Hide the Verdict, Keep the Evidence: Inquiry Dialogue for Multi-Agent LLM Debate

Abstract

Multi-Agent Debate (MAD) typically requires agents to publicly disclose their initial answers to peers before any argument is exchanged, creating a persuasion-like setting vulnerable to conformity. We introduce INQUIRY, a verdict-blind, inquiry-inspired protocol in which agents still form a private answer each round, but explicit verdicts are withheld from peers until a final decision round. Across four multiple-choice QA benchmarks, we compare INQUIRY against Single-Agent, Independent (no-debate) aggregation, MAD-SoM, and MAD-Conformist baselines using both a frontier and a small open-weight (SLM) judge panel. On the SLM panel, INQUIRY achieves the highest multi-agent accuracy on every dataset, outperforming Independent by 5.7-20.6 percentage points, while SoM, the dominant early-commitment protocol, is often the weakest multi-agent method; only on CommonsenseQA does INQUIRY fall short of the single strongest individual judge's high solo accuracy. On the frontier panel, near-ceiling accuracies leave far less room for protocols to differ, suggesting verdict exposure matters most for weaker judges. Correct Reasoning Survival and Minority Truth Recovery diagnostics show that INQUIRY better preserves correct majorities and recovers substantially more minority-held correct answers than the other baselines. A further ablation finds that a mandatory argument-structure schema (Claim, Grounds, Warrant) harms small-panel accuracy on three of four datasets but has no detectable effect at frontier scale. These results identify verdict exposure as a consequential MAD design choice whose effects depend on model capability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.