Strategic Competition and Convergence in LLM Leaderboards
Abstract
Leaderboards are a popular way to compare large language model (LLM) capabilities and can influence how model developers invest their resources. We ask when competition on head-to-head leaderboards leads developers to produce similar models, which we call monoculture, or models specialized in different capabilities. We study this question by formulating leaderboard competition as a game in which model developers invest resources across capabilities to maximize their win rate. We show that similar production functions can induce monoculture, while comparative advantage can sustain specialization. Turning to how competition evolves over time, we give conditions under which diminishing returns and model catch-up leads to crowding at the frontier. Using data from LM Arena and OpenRouter, we find evidence of both competitive catch-up and persistent differentiation: developers catch up where they lag while retaining category-specific advantages, alongside an increasingly crowded frontier and dispersed model usage. Together, our results characterize conditions under which competitive investment can generate convergence while preserving differences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.