Route Teams, Not Agents: Team-Value Learning for Multi-Agent Reasoning
Abstract
Effective multi-agent reasoning requires selecting agents that work well together rather than simply assembling strong individuals. Complementarity, redundancy, and team-size effects make individual suitability an incomplete guide to collective performance. We introduce TeamRoute, which separates offline outcome learning from online constraint-aware routing. Its Experience Brain combines context-dependent role, pairwise, and cardinality factors in a latent score space and decodes them into quality and resource estimates. The Rational Scheduler uses skill- and capacity-aware optimal transport to generate candidates, then applies the frozen predictor to screen resource budgets and minimize predicted cost within a quality tolerance. This separation allows runtime budgets and preferences to change without retraining the predictor. Our analysis characterizes a limitation of additive latent scores and derives a conditional quality-loss bound involving prediction error, quality tolerance, and candidate restriction. Experiments across diverse reasoning benchmarks show improved average accuracy and favorable accuracy–budget trade-offs. Ablations and diagnostic experiments further support structured set-level prediction, selection among unseen agent combinations, and transport-guided candidate generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.