Strategic Diversity Does Not Scale: The Bounded Geometry of LLM Collective Behavior
Abstract
Multi-model AI systems are often motivated by an expectation that independently developed language models provide complementary judgments. We examine this expectation using 50 LLMs from 26 model families evaluated on 20 strategic-game conditions spanning signaling, coordination, Bayesian reasoning, and deception. We represent model responses using text embeddings and study the geometry of both individual responses and model-level behavioral signatures. As the evaluated ecosystem grows, we observe that dimensionality of the pooled response space increases with sharply diminishing marginal returns. At 50 models, principal component analysis (PCA) over model-level signatures requires 16 components to explain 90% of between-model variance. Model signatures also exhibit high pairwise similarity and substantial cross-family geometric alignment. These patterns persist in action-only representations, alternative dimensionality estimators, and an open-ended generation dataset. Together, our results indicate that nominal growth in models and providers produces less behavioral expansion than model count alone would suggest. We further investigate strategic diversity when models are aggregated following three patterns: majority vote, multi-agent discussion, cross provider composition. Majority voting reinforces model's dominant responses. Multi-agent discussion among diverse models produces no clear gain in strategic score, while critique from another provider often changes the decision but does not reliably improve it. Effective collective systems thus require measured coverage of useful behaviors, not simply more models, providers, or calls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.