acceptodds
Under review as a conference paper at ICLR 2027

Understanding LLM Multi-Agent Scaling: Agent Composition and Evidence Coverage

Abstract

Collective intelligence depends not only on how many agents participate, but also on what each contributes. We examine how model and persona composition shapes the returns from scaling LLM-based multi-agent systems. An evidence-coverage framework shows that equal individual coverage need not yield equal collective recovery: observations can overlap in what they miss. Across seven benchmarks and groups of two to sixteen agents, combining model and persona diversity achieves the highest seven-task macro-average accuracy among the tested configurations at every evaluated agent count in both voting and debate. Persona diversity provides larger average gains in mixed-model than in single-model groups across these settings, although benefits vary across tasks. We further analyze multi-round response trajectories using semantic effective rank to characterize how output diversity varies across agent configurations. Together, these findings highlight agent composition as a key dimension of LLM multi-agent scaling, motivating a focus on the complementary contributions of additional agents rather than their number alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.