You Don't Need the Whole Meta-Dataset: Budgeted Model Evaluation on Unseen and Unlabeled Data
Abstract
Predicting model performance on an unseen and unlabeled workload commonly starts by measuring its distribution shift from the training data. To convert this shift into an accuracy estimate, prior works train an evaluator on a diverse meta-dataset, using the shift between each subset of the meta-dataset and the training data as input, and the model accuracy on that subset as the label. However, obtaining these labels is expensive because each model must be run and evaluated on every subset. The more subsets the meta-dataset contains, the more forward passes are required from each reference model to obtain its accuracy labels before the evaluator can be trained. To overcome this limitation, we present ActiveEvaluator, which trains the same evaluator on a small representative subset of the meta-dataset instead of the full one, reducing the cost and latency of model evaluation to nearly one third while matching or even outperforming full-set training. ActiveEvaluator measures subset similarity using Hausdorff distance and maximizes a facility-location objective that assigns each subset to its closest representative, giving redundant subsets little marginal gain. The objective is monotone and submodular, so greedy selection comes with a provable performance guarantee. We validate ActiveEvaluator across Text2SQL, image classification, and node classification, showing that the same budgeted selection principle generalizes across structurally different tasks. As data and model pools grow faster than the resources available to evaluate them, this opens a direction toward fast, budgeted evaluation of new models on unlabeled data. It also makes evaluation more practical in domains where data are limited. Our code is available at https://anonymous.4open.science/r/ActiveEvaluator-00B1/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.