AdaEval: Adaptive Importance Weighting for Reliable Model Evaluation under Unknown Covariate Shift
Abstract
Reliable model evaluation in the wild often relies on the AUTOEVAL framework, which augments a small human-labeled dataset with massive synthetic labels to produce low-variance performance estimates. A critical limitation is that AUTOEVAL assumes either that the labeled and unlabeled data share the same distribution or that the true importance weights are known, both of which are unrealistic under real-world covariate shift. We propose ADAEVAL, a method that learns importance weights directly from unlabeled data by training a discriminative classifier to distinguish labeled from target inputs. The resulting weighted estimator re-targets the labeled sample to the shifted distribution without any prior knowledge of the shift. We prove consistency and asymptotic normality under mild conditions on the weight estimation error. On ImageNet with simulated covariate shift, standard AUTOEVAL collapses to a coverage of just with labeled points; ADAEVAL achieves – coverage against PPI++'s –, and reduces MSE by – relative to classical estimation. On ProteinGym, under a strong fitness-biased shift, ADAEVAL reduces absolute MSE bias by – while coverage decays to near zero as grows large, a residual gap we characterize explicitly. ADAEVAL reduces bias and MSE under unknown covariate shift, with a fully characterized residual coverage gap; it is not a general solution for statistically valid confidence intervals across all regimes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.