acceptodds
Under review as a conference paper at ICLR 2027

Efficient Evaluation of Reasoning Benchmarks with Limited Historical Information

Abstract

Accuracy estimates on reasoning benchmarks can be noisy, with uncertainty arising from the small number of benchmark items and the stochastic generation of responses by language models. Repeated evaluation can reduce this uncertainty but requires additional generations, raising the question of how to allocate a generation budget across items to minimize error in estimating benchmark accuracy. We formulate this problem as sequential item selection and develop adaptive sampling policies that leverage intermediate evaluation results from the target model to guide subsequent sampling. Our sampling policies also incorporate structural information about item difficulty from historical evaluations of other models on the benchmark, such as difficulty strata or a relative difficulty order. We provide theoretical analysis showing that informative difficulty strata can reduce estimation error, while a reliable relative difficulty order can tighten an estimation error bound. Experiments show that the resulting policies reduce error in estimating benchmark accuracy across a range of budgets and levels of available historical information. We summarize practical guidelines for choosing an evaluation strategy according to the historical information available, enabling efficient and reliable evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.