Protecting the Group, Not Just Its Members: Feasible Differential Privacy for Groups
Abstract
Large amounts of data are required to infer robust statistics or train large machine-learned models. Not only a public release of such data can reveal sensitive information of the individuals, but such information can also be inferred from a statistical estimator or learned predictor. Differential privacy (DP) provides strong guarantees for protecting individual records; however, these guarantees do not prevent sensitive properties of groups from being inferred. In many applications including healthcare, protecting the characteristics of a subpopulation within an institution can be as important as protecting individual records, for example disease prevalence within a demographic group at a hospital. We introduce *Differentially Private Attributes of Groups (DPAG)*, a framework that treats such groups as the atomic units of privacy. Rather than obtaining group protection by scaling an individual-level privacy guarantee with group size, DPAG directly bounds and perturbs each protected group's contribution. We develop mechanisms for centralized learning and extend the framework to cross-silo federated learning that is regularly used in healthcare scenarios to transfer institutional knowledge for machine learning without sharing the underlying data. We evaluate DPAG on centralized and federated clinical data, comparing direct group-level protection with standard DP mechanisms and conventional group-privacy approaches. Our results show that **DPAG can make group-level protection feasible without a privacy cost that grows with the number of individuals in the group, while retaining useful model utility**.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.