Learning What to Debate: Adaptive Argument Scheduling for Multi-Agent Reasoning
Abstract
Multi-agent debate (MAD) lets Large Language Models (LLMs) read one another's responses and iteratively revise their reasoning. Yet MAD is judged by collective decision quality, while computation is allocated through coarse, full-response rounds. This can repeatedly revisit low-value reasoning, waste computation, and propagate erroneous arguments. We instead formulate MAD as selective computation over a **public argument state**, organizing independently generated rationales into a shared representation of claims, candidate decisions, evidence links, and support and conflict relations. Our method predicts the rescue and damage of candidate revisions, ranks focused revision actions by predicted rescue and call cost, and stops through a damage-aware gate. Opening-relative **promotion**, **acquisition**, and **damage** distinguish revision gains from the starting ensemble. Across eight reasoning benchmarks, AAS-MAD achieves 70.76% pooled accuracy, exceeding the strongest evaluated debate baseline by 1.03 percentage points with 28.84–44.34% fewer tokens. Measured from its own opening rather than from the best baseline, its 1.45-percentage-point gain is the largest among the five MAD methods; at the evaluated policy endpoints, all successful corrections promote existing candidates. On our curated FIFA 2026 dataset, AAS-MAD reaches 68.59% accuracy, improving on the strongest evaluated debate baseline by 1.92 percentage points. These results support allocating deliberation to arguments that can still improve the shared decision, limiting expenditure on unproductive revisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.