acceptodds
Under review as a conference paper at ICLR 2027

Route Teams, Not Agents: Team-Value Learning for Multi-Agent Reasoning

Abstract

Effective multi-agent reasoning requires selecting agents that work well together rather than simply assembling strong individuals. Complementarity, redundancy, and team-size effects make individual suitability an incomplete guide to collective performance. We introduce TeamRoute, which separates offline outcome learning from online constraint-aware routing. Its Experience Brain combines context-dependent role, pairwise, and cardinality factors in a latent score space and decodes them into quality and resource estimates. The Rational Scheduler uses skill- and capacity-aware optimal transport to generate candidates, then applies the frozen predictor to screen resource budgets and minimize predicted cost within a quality tolerance. This separation allows runtime budgets and preferences to change without retraining the predictor. Our analysis characterizes a limitation of additive latent scores and derives a conditional quality-loss bound involving prediction error, quality tolerance, and candidate restriction. Experiments across diverse reasoning benchmarks show improved average accuracy and favorable accuracy–budget trade-offs. Ablations and diagnostic experiments further support structured set-level prediction, selection among unseen agent combinations, and transport-guided candidate generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.