xHetMed-FL: Explainable Federated Shared Representation Learning Across Heterogeneous Medical Imaging Modalities
Abstract
Medical images are usually stored in different hospitals and with different imaging modalities and are often subject to different storage formats, privacy concerns, modality-specific characteristics and different security restrictions, which hinders centralized learning. Current federated learning methods usually assume a uniform class or learning goal across clients and the raw patient data cannot be shared. To address these challenges, we suggest a multimodal federated learning scheme, xHetMed-FL, using 2D/3D CNNs to create a shared 128-D latent representation while federating only a few projection parameters. The proposed framework is organized around four heterogeneous clients representing SPECT, Chest X-ray, MRI, and CT+PET imaging. Rather than forcing these modalities into a common input architecture, each client retains a modality-specific preprocessing pipeline and encoder. The SPECT client uses verified pathological and non-pathological labels for supervised disease classification, whereas the Chest X-ray, MRI, and CT+PET clients learn representations without disease labels through a cosine-similarity-based self-supervised objective. Only shared projection is communicated and aggregated by the federated server, while the modality-specific encoders, client-specific components, prediction heads, and Batch Normalization parameters remain local. To address optimization under heterogeneous client conditions, FedProx is incorporated into local training, providing proximal regularization with respect to the global shared parameters, while FedBN preserves client-specific Batch Normalization statistics. The obtained SPECT client's validation accuracy, F1 score, and ROC-AUC were 82.35%, 82.86%, and 91.77%, respectively. The cosine similarity in the Chest X-ray, MRI, and CT+PET clients was found to be 1.0000, 1.0000, and 0.9992, respectively, while the variance in their latent representations was 0.3185, 0.3168, and 0.2965, respectively. Grad-CAM is used to interpret the trained network at the post-training stage, providing modality-specific explanations while preserving the heterogeneous local components. Overall, the proposed framework shows the viability of learning shared representations from heterogeneous medical imaging modalities with different learning goals while keeping modality-specific components and keeping the models interpretable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.