acceptodds
Under review as a conference paper at ICLR 2027

Sample Efficient Best-LLM Identification via Active Pairwise Preference Elicitation

Abstract

Practitioners deploying Large Language Models (LLMs) routinely face the following problem: among the many candidate LLMs available, which one performs best on my task? Public leaderboards rank models on generic benchmarks that may not reflect a particular use case, while building a custom evaluation benchmark in-house is often prohibitively expensive. In this paper, we study the problem of best-LLM identification when the practitioner has access to a small unlabeled pool of prompts drawn from their use case. We elicit from an annotator pairwise preferences between two candidate responses to the same prompt, and we ask how to select the next triple to label so that the best model is identified with as few labels as possible. We introduce TOPS (Top-focused Online Preference Sampling), an active acquisition strategy that maintains a Bradley-Terry posterior over model abilities and samples comparisons using acquisition scores that combine response dissimilarity with estimated reductions in uncertainty about the margins between the leader and its challengers. Across 17 settings, TOPS achieves the best average rank for cumulative regret, final regret, and the rank of the selected model. Using 500 queries, TOPS is able to identify the single best model from 20 candidates in 62% of runs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.