Efficient Evaluation of Reasoning Benchmarks with Limited Historical Information
Abstract
Accuracy estimates on reasoning benchmarks can be noisy, with uncertainty arising from the small number of benchmark items and the stochastic generation of responses by language models. Repeated evaluation can reduce this uncertainty but requires additional generations, raising the question of how to allocate a generation budget across items to minimize error in estimating benchmark accuracy. We formulate this problem as sequential item selection and develop adaptive sampling policies that leverage intermediate evaluation results from the target model to guide subsequent sampling. Our sampling policies also incorporate structural information about item difficulty from historical evaluations of other models on the benchmark, such as difficulty strata or a relative difficulty order. We provide theoretical analysis showing that informative difficulty strata can reduce estimation error, while a reliable relative difficulty order can tighten an estimation error bound. Experiments show that the resulting policies reduce error in estimating benchmark accuracy across a range of budgets and levels of available historical information. We summarize practical guidelines for choosing an evaluation strategy according to the historical information available, enabling efficient and reliable evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.