Which Model Should Be Chosen? Joint Performance Estimation and Ranking of Black-Box Models on Unlabeled Data
Abstract
Evaluating a pool of models on an unseen and unlabeled dataset requires both accurate performance estimates and reliable rankings. Many existing evaluation methods estimate performance for one model at a time, often requiring model-specific training or adaptation, access to training data, or internal representations to characterize distribution shifts. These requirements limit evaluation of black-box models. Moreover, methods that evaluate models independently miss cross-model evidence that could improve ranking within a model pool. We propose PoolEvaluator, a framework for jointly estimating and ranking model performance without target labels or access to model internals. It estimates prior accuracies from relevant calibration datasets, then combines these priors with model agreement through expectation-maximization to infer latent correctness and refine each model's accuracy estimate using a closed-form update. An external budgeted judge is used later to refine ambiguous latent answers. Experiments across Text2SQL, image classification, and node classification show improved accuracy estimation and ranking with lower end-to-end evaluation latency than existing methods. This opens a new direction for jointly evaluating and ranking black-box model pools on unlabeled data through shared prediction evidence. Our code is available at https://anonymous.4open.science/r/PoolEvaluator-664E.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.