Self-Improvement of Language Models by Post-Training on Multi-Agent Debate
Abstract
Self-improvement, where models improve beyond their current performance without external supervision, remains a challenge. The core difficulty is sourcing a training signal stronger than what the model itself can currently produce. Majority voting has been shown to provide such a signal by aggregating over multiple samples, helping mitigate some of the inconsistencies in LM reasoning. In this work, we study multi-agent debate—where models collaborate and exchange reasoning over multiple rounds—as an alternative to single-round majority voting for constructing self-improvement signals. We introduce Multi-Agent Consensus Alignment (MACA), which uses reinforcement learning (RL) to post-train models on debate-derived signals. We study scalar consensus rewards and preference learning over majority and minority reasoning traces, comparing them with SFT on consensus-supporting traces. MACA improves (1) multi-agent debate accuracy (up to +26.87% on MATH), (2) single-agent accuracy (up to +21.51% on MathQA), and (3) self-consistency (up to +27.6% on GSM8K). We also observe gains on unseen benchmarks, including +16.3% on GPQA and +11.6% on CommonsenseQA. In a separate rollout-matched study, we evaluate online self-debate GRPO, in which one policy debates among its own rollouts during RL. The resulting policy improves held-out GPQA accuracy by 3.39% over the zero-shot base model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.