acceptodds
Under review as a conference paper at ICLR 2027

Self-Improvement of Language Models by Post-Training on Multi-Agent Debate

Abstract

Self-improvement, where models improve beyond their current performance without external supervision, remains a challenge. The core difficulty is sourcing a training signal stronger than what the model itself can currently produce. Majority voting has been shown to provide such a signal by aggregating over multiple samples, helping mitigate some of the inconsistencies in LM reasoning. In this work, we study multi-agent debate—where models collaborate and exchange reasoning over multiple rounds—as an alternative to single-round majority voting for constructing self-improvement signals. We introduce Multi-Agent Consensus Alignment (MACA), which uses reinforcement learning (RL) to post-train models on debate-derived signals. We study scalar consensus rewards and preference learning over majority and minority reasoning traces, comparing them with SFT on consensus-supporting traces. MACA improves (1) multi-agent debate accuracy (up to +26.87% on MATH), (2) single-agent accuracy (up to +21.51% on MathQA), and (3) self-consistency (up to +27.6% on GSM8K). We also observe gains on unseen benchmarks, including +16.3% on GPQA and +11.6% on CommonsenseQA. In a separate rollout-matched study, we evaluate online self-debate GRPO, in which one policy debates among its own rollouts during RL. The resulting policy improves held-out GPQA accuracy by 3.39% over the zero-shot base model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.