acceptodds
Under review as a conference paper at ICLR 2027

Stop, Continue, or Escalate: Adaptive Control of Multi-Agent Debate

Abstract

Multi-agent debate can improve the reasoning of compact language models, but applying the same amount of debate to every query is inefficient. Easy queries are often resolved within one or two rounds, some disagreements benefit from further debate, and others persist or converge to shared errors. When local debate remains unreliable, the query can be escalated to a stronger external model, although access to it may be limited by cost or usage constraints. This raises a practical question of whether, after each round of local debate, the system should stop, continue debating, or escalate the query. We formulate this problem as budget-aware, round-by-round control of multi-agent debate. Our framework learns to resolve the current disagreement from the evolving debate trajectory and chooses among the three actions by estimating the value of the current local answer, further local debate, and escalation, while an online mechanism keeps escalation within a prescribed budget. Across four reasoning benchmarks and four agent configurations (three homogeneous and one heterogeneous), our method achieves a better accuracy-computation trade-off than fixed-depth debate as well as routing, cascading, and adaptive-control baselines. At a 10% escalation rate, it improves accuracy by 2.3 percentage points over the strongest baseline while using 18% fewer local model calls than full-length debate. Further analysis shows that our answer resolver improves answer selection under disagreement, whereas the debate controller avoids unnecessary debate rounds through selective continuation and prioritizes escalation when predicted local outcomes are weak.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.