Feature Attribution from Multi-Source Data with Heterogeneous Observation Patterns
Abstract
In multi-source learning, different sites typically observe different feature subsets, so pooling data across sites produces blockwise missingness and distributional shift. Shapley-based feature attribution has become a standard tool for model interpretability, yet existing estimators require fully observed evaluation data. We propose DASH (Data-fused Attribution via Shapley), which derives the influence function of the constrained weighted-least-squares Shapley estimator and uses it to construct site-specific control variates from partially observed auxiliary sites, reducing variance without imputing missing features. A permutation-based screening step excludes incompatible sources, and data-adaptive calibration optimally weights each site. In simulations, DASH achieves – lower MSE than the single-site estimator and – lower error than imputation baselines, with screening power at moderate misalignment. On environmental monitoring and multi-center clinical data, DASH reduces error by – over the single-site baseline, while imputation can increase error in the clinical setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.