acceptodds
Under review as a conference paper at ICLR 2027

Sample Efficient Ranking of Language Models Using Pairwise Comparisons

Abstract

The frontier of large language model (LLM) research is advancing rapidly, with new models being released every few weeks. In this fast-paced environment, rigorous evaluation is crucial to accurately assess a model's quality. Yet evaluation prompts vary in how much evidence they provide about the relative quality of a model pair. In this work, we study how to reduce evaluation cost by adaptively selecting both model pairs and evaluation prompts while controlling comparison errors. Our contribution is twofold: First, we formulate pairwise LLM evaluation as a sequential hypothesis-testing problem and develop an e-process-based procedure. We further derive a sampling distribution for actively selecting evaluation prompts. Second, we construct a complete LLM ranking by using this pairwise evaluation procedure as a black-box comparison function within a sorting algorithm. To further reduce evaluation cost, we adaptively allocate the error budget across pairwise comparisons. Empirically, across both synthetic and real-world setups, we observe that our approach can efficiently rank LLMs using only 7-22% of the total evaluation budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.