BayesAME: Bayesian Active Model Evaluation
Abstract
Evaluating large generative models across benchmarks is time-consuming and costly, motivating methods that can estimate full benchmark performance by evaluating models on only a subset of items (coreset). While current literature mostly requires the practitioner to input a coreset size, we introduce BayesAME, a sequential active Bayesian framework that uses historical data and an information-gain criterion to dynamically select items to include in the coreset thus allowing automatic determination of its size. Additionally, we show how a multi-target extension further reduces coreset size by leveraging model correlations. Extensive experiments demonstrate that BayesAME consistently outperforms existing methods. Crucially, our analysis resolves recent skepticism in the literature by showing that a) non-random coreset selection is advantageous over random selection and b) one can reliably improve upon the random sample mean baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.