Having It Both Ways: Personalization and Generalization in Federated EHR Generation
Abstract
Patient privacy regulations restrict the sharing of electronic health records (EHRs), limiting their use for clinical machine learning. While sharing realistic synthetic data can bypass these legal restrictions, EHR generation approaches trained on a single hospital's data generalize poorly to patient populations elsewhere, and their quality is often limited by the quantity of patient data available. To address these issues of data sparsity, federated learning enables training on data across many hospitals by only sharing model updates, allowing hospitals to learn statistical features of other hospital populations without leaking protected health information (PHI). In this paper, we first identify a trade-off inherent to this setting between generalization and personalization: joint training improves transfer across hospitals but degrades performance on each hospital's local distribution. We characterize this trade-off across a range of federated methods and propose FedLIBRA, an approach that directly pushes the frontier of personalization–generalization, enabling the generation of high-quality patient data for any hospital.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.