Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate
Abstract
The reasoning abilities of large language models (LLMs) have been substantially improved by reinforcement learning with verifiable rewards (RLVR). At test time, collaborative reasoning through Multi-Agent Debate (MAD) has emerged as a promising approach for enhancing LLM performance. However, RLVR typically trains isolated solvers, while multi-agent RL often optimizes MAD as a system. We therefore argue that improving MAD first requires strengthening each participant's ability to evaluate competing reasoning and decide when to revise or retain its answer. We show that reassessment improves average correctness when corrections of wrong answers outweigh changes from correct to incorrect answers, and that maximizing correctness rewards also maximizes this net accuracy gain. Guided by this analysis, we propose Self-Debate Reinforcement Learning (SDRL), which post-trains a single LLM to solve problems independently and reason over peer responses during multi-agent debate. SDRL builds debate contexts from sampled solutions, with frequency-based pairing prioritizing the most prevalent answer disagreements, and generates responses conditioned on these contexts. Finally, SDRL jointly optimizes both the initial and debate-conditioned responses, yielding a model that is effective as both a standalone solver and a debate participant. Experiments across multiple base models and reasoning benchmarks show that SDRL strengthens single-model reasoning and improves MAD performance across diverse debate protocols and agent configurations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.