ARMAD: Adaptive Routing for Heterogeneous Multi-Agent Debate in Large Language Model Reasoning
Abstract
Multi-agent debate (MAD) can improve large language model reasoning, but fixed-depth protocols allocate the same interaction budget to every input, potentially wasting calls when independent answers already agree and risking harmful over-deliberation. We introduce ARMAD, a training-free, black-box-compatible controller that adapts both debate initiation and duration within a fixed heterogeneous model pool. After collecting independent responses, Pre-debate Agreement Routing (PAR) bypasses debate when agreement among rule-normalized final answers reaches a predefined routing threshold. For routed inputs, the Early Agreement Stopping Evaluator (EASE) recomputes the same agreement score after each synchronized revision round and terminates interaction when a separate stopping condition is satisfied or the maximum depth is reached. Responses from the selected round are then aggregated using a deterministic voting rule. We evaluate ARMAD on five reasoning benchmarks against same-pool fixed-depth controls and call-matched routing controls. ARMAD improves over Hetero-Vote (Round-0 voting) on every benchmark and achieves 82.7% macro-average accuracy with 3.96 model calls per item, 67.0% fewer than same-pool fixed three-round debate. At the same call budget, agreement-guided routing achieves higher accuracy than random and response-length allocation on all five benchmarks. Shallow fixed-round policies achieve modestly higher macro-average accuracy at larger call budgets, revealing an accuracy–call trade-off. Together, these results establish answer agreement as an effective control signal for adaptive, call-efficient heterogeneous debate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.