When Is Further Collaboration Worthwhile? Cost-Aware Multi-Agent Reasoning via Counterfactual Branching
Abstract
Multi-Agent Debate (MAD) can enhance the reasoning capabilities of Large Language Models (LLMs), but additional collaboration rounds introduce substantial token overhead. Existing approaches typically make decisions at the instance level to determine whether collaboration should be initiated. However, they do not continuously reassess the need for further interaction as the reasoning state evolves. Consequently, they could incur redundant interactions on tasks that have already been resolved while providing insufficient refinement for those that still require further reasoning. To address this limitation, we introduce RLMAS, a cost-sensitive reinforcement learning framework for adaptive multi-agent reasoning. RLMAS augments the collaboration policy state with capability-aware confidence signals and trains a state-dependent policy using a reward that balances improvements in answer quality against incremental token costs. We introduce an offline counterfactual branching scheme to improve per-state action coverage beyond single-path trajectories by expanding alternative actions from shared reasoning states. The resulting feedback for different actions helps the policy learn when to stop and which refinement action to select. During deployment, the learned collaboration policy selects actions based on the current reasoning state without constructing counterfactual branches. Across diverse reasoning benchmarks and model scales, RLMAS reduces token usage by up to 92.6% while improving accuracy by up to 8.8% relative to baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.