Privacy Considerations in Multi-Institutional Data Harmonization Through a Category-Theoretic Lens
Abstract
Institutions pool data to overcome local scarcity, and filtering-based matching harmonizes data pooling by admitting only samples that align with a reference site. We show it is unsafe for privacy, and that the two failures have a common cause. For an institution/site whose population not harmonized with respect to other institutions, the samples that align best with the reference institution/site are precisely the atypical ones in its own distribution, so the attempt to harmonize data forces each site to contribute an unrepresentative subset of itself. What leaks is then not a patient record but the distributional identity of the institution. We therefore propose pooling specific a – DP metric () for data pooling, a task for which privacy has not been studied and is fundamentally different from existing DP setups. Using category theory helps us demonstrate and understand the natural tension between data harmonization and , proving a structural incompatibility where improving data harmonization (utility) strictly degrades (privacy) for confounded institutions under data distribution shifts. To resolve this, we propose privacy-budgeted minimax matching (PBMM) that jointly optimizes for distributional alignment, sample efficiency, and differential privacy. Across synthetic, multi-site fMRI (homogeneous institutional distribution) and six-modality medical imaging benchmarks (heterogeneous institutional distribution), PBMM preserves the harmonization of filtering while cutting peak in-domain exposure by over and recovering samples that filtering discards.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.