High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning
Abstract
To commit to buying external data or participating in a collaborative learning process, one must decide whether the additional data will improve prediction enough to justify the cost. This comes with several challenges: (i) in many settings, e.g. in healthcare, the decision relies only on aggregated statistics available publicly, rather than individual-level data; (ii) covariate and model shifts can induce negative transfer, so the additional data deteriorates rather than improves performance; (iii) if the data is sensitive, its privatization requires the injection of noise, which can also offset the benefit of a larger sample size. In this paper, we model the problem of dataset selection through high-dimensional linear regression with multiple heterogeneous sources and a weighted ridge estimator. Our approach uses only summary statistics and it gives privacy guarantees either on labels only or jointly on features and labels, in terms of -zero-concentrated differential privacy. The main technical contribution is a deterministic equivalent of the test error, which captures the interactions between sample size, covariance structure, model shift, regularization and privacy noise. Our theory allows to optimize hyper-parameters (weights and ridge regularizers) and, more broadly, to decide when private external datasets are useful. Overall, our work provides a theoretically tractable foundation for private transfer learning, and we support this foundation via experiments on both synthetic and real-world datasets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.