Learning under Mechanism-Induced Covariate Shift: Finite-Scale Coverage and Source Allocation
Abstract
Simulators and physical acquisition systems can generate abundant data yet still leave deployment-relevant neighborhoods poorly observed, even when those neighborhoods are reachable. We study this as mechanism-induced covariate shift: the conditional law is shared, but a generator's acquisition law and forward map determine where training examples appear in task space. The key quantity is the finite-scale local-mass field—the probability that one generated example falls within a neighborhood of each deployment point. This quantity connects the generator directly to finite-budget learning. For Nadaraya–Watson (NW) regression, it yields a finite-sample risk certificate with constants uniform over the generator family, and on a controlled one-dimensional common-support family the optimized certificate matches minimax risk up to family-uniform constants. Keeping the full spatial field matters: generators can have identical normalized-Jacobian distributions and identical distributions of source-density values at deployment draws while having different finite-scale coverage. The same field also turns upstream data collection into a pre-training decision problem: task matching removes generator-dependent risk differences at fixed labels, exact rejection sampling exposes the corresponding proposal cost, and source mixing gives a convex allocation problem that can exploit spatial complementarity. On a controlled 23-design family, the coverage score tracks exact finite- NW risk with median target-wise Spearman correlation 0.97–1.00 across six settings. Estimated from unlabeled forward samples, finite-scale coverage can therefore guide source selection and allocation before training, without learner-side feedback.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.