acceptodds
Under review as a conference paper at ICLR 2027

Active Testing of Large Language Models via Approximate Neyman Allocation

Abstract

Large language models (LLMs) are showing superb performance from general-QA to agentic tasks, mathematical frontier and even self-improvement in niche areas. Evaluating these models on domain-specific tasks often requires expert annotation, making it costly to label the entire evaluation pool. Active testing aims to alleviate this bottleneck by approximating the evaluation result from a small but informative subset of the evaluation pool. Prior active-testing approaches are either built for traditional deep learning models or restrict LLM answers to a single token. We remove this restriction and introduce a training-free active testing algorithm for modern reasoning LLMs on tasks with objectively verifiable answers. Our method leverages semantic entropy from surrogate models to stratify the evaluation pool and then conducts approximate Neyman allocation based on signals extracted from these surrogates. Across multiple language and multimodal scenarios, several surrogate-target pairs and different label budgets, our method attains an average 87.0% relative MSE against Uniform Sampling, translating to 15.0%–28.0% label savings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.