StageMAD: Process-Controlled Communication for Multi-Agent Debate
Abstract
Performing multi-step reasoning with large language models based on a single sampled trajectory remains vulnerable to local errors in intermediate steps, as errors may propagate through subsequent reasoning and finally lead to a wrong answer. To mitigate this limitation, multi-agent debate (MAD) samples multiple reasoning trajectories and allows agents to exchange information and revise their reasoning before answer aggregation, which is typically performed through majority or plurality voting. However, most existing MAD methods lack reliability constraints on both the content and propagation of exchanged information, so incorrect reasoning can also spread across agents like correct reasoning. As individual responses are not reliable, unverified erroneous reasoning may influence other agents, suppress an initially correct minority, and even produce an incorrect consensus. To address this problem, we propose StageMAD, a process-guided multi-agent debate framework consisting of three stages: local self-correction, process-evidence-based directed logical guidance, and mutual critique over residual disagreements. By localizing questionable reasoning steps and selectively controlling information propagation, StageMAD enables reliable reasoning to gain broader support while limiting the spread and accumulation of incorrect answers. Experiments on mathematical reasoning and language understanding datasets demonstrate the effectiveness of our proposed framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.