Adaptive Multi-Source Data Collection under Distribution Shifts
Abstract
When labeled data from a target population are scarce, one can acquire additional training observations from related sites. These sites often differ in their covariate distributions and conditional outcome distributions , making their value for target prediction unequal and initially unknown. We present an adaptive framework for planning batched data collection across source sites under a fixed budget. We formulate the problem as a Markov decision process, in which states are posterior beliefs about source response functions and actions are site selection for the next batch. Our objective is to reduce disagreement among predictors fitted to the available target labels, measured through their prediction variance over the target population. We focus on one-step lookahead, which selects the next batch to minimize the expected remaining predictor variance. In a tractable Gaussian model, we show that exact one-step lookahead is asymptotically optimal for this objective under regularity and sufficient-sampling conditions. Our framework supports a range of posterior modules (e.g., BNNs and TabPFN) for simulating source labels, as well as different prediction models. Experiments with linear and neural synthetic data and U.S. census income records show improved target prediction over several data collection baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.