acceptodds
Under review as a conference paper at ICLR 2027

When Is a Team Worth Convening? Separating Aggregation, Sampling, and Member Selection

Abstract

Multi-agent systems (MAS) increasingly convene multiple heterogeneous language models as a team, and the resulting performance gains are often attributed to collaboration. A team's inference, however, bundles collaboration with other mechanisms, including how candidate answers are aggregated, how much sampling budget is spent, and how the single-member alternative is selected, so an observed team gain cannot by itself identify collaboration as the source of the improvement. In this paper, we isolate these sources through controlled experiments on verifiable tasks with open non-thinking models of up to 14B parameters. We use two simple voting policies throughout, named after the answer camps (groups of responses giving the same answer) they count: CAMP-1 takes a single plurality vote, whereas CAMP adds votes only when first-pass agreement is low. Our findings show that the apparent team advantage changes with each of the three controls. The relative ranking of voting and synthesis depends on the aggregator, and with full candidate inputs, synthesis shows no average loss to an order-neutral vote. At a comparable budget, a hindsight-selected single member outperforms the tested voting teams, whereas neither side has an established advantage when selection is separated from evaluation. The relative value of routing depends strongly on label budget and pool composition: teams outperform the router at small label budgets in the development-selected pool, but this advantage is not established at larger budgets, while broader pools with a dominant member favor routing. Together, these results provide a practical decomposition for deciding whether to convene a team: before attributing a team gain to collaboration, one should determine how much comes from the aggregation rule, additional sampling, and the information available for member selection. This reframes a basic question for MAS—not only how a team should collaborate, but whether the task warrants a team at all.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.