acceptodds
Under review as a conference paper at ICLR 2027

Sample-Efficient LLM Evaluation via Pure Exploration under Task Heterogeneity

Abstract

We study sample-efficient LLM evaluation from preference feedback under task heterogeneity, where LLM comparisons are actively collected across diverse tasks to identify the best-performing model overall. We formulate this problem as fixed-confidence pure exploration under a contextual multinomial logit model with task-dependent model features, inducing heterogeneous preferences coupled through a shared latent parameter. Using a change-of-measure argument and a local quadratic approximation, we derive a Fisher-geometric separation criterion that characterizes the identification difficulty. A tractable relaxation of this criterion admits an explicit directional-variance representation, suggesting that comparisons should be allocated to reduce the worst uncertainty over decision-changing alternatives. Building on this principle, we develop Multi-FW, a global Frank–Wolfe procedure that coordinates task-conditioned comparison allocations. Theoretically, combining estimator concentration and consistency with our proposed task-aware tracking, we establish asymptotically optimal separation growth despite exogenous task arrivals. For any fixed , suitably tuning forced exploration yields an sample-complexity upper bound for Multi-FW, approaching the lower bound's logarithmic confidence dependence. Finally, we evaluate Multi-FW in a semi-real environment constructed from the Arena Human Preference 140K dataset, using logged human preferences and real LLM responses. Multi-FW achieves faster separation growth than task-agnostic Elo-style, uniform-random, and dueling-style baselines, with at least a -percentage-point gain in trajectory-averaged identification accuracy and more than a reduction in the comparison budget needed for stable accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.