FedMAGE: Architecture-Aware Graph Hypernetwork for Heterogeneous Multimodal Federated Learning
Abstract
Multimodal federated learning aims to collaboratively train multimodal models across decentralized clients while preserving data privacy. However, existing methods struggle with highly heterogeneous client tasks and architectures, often rely on public data for knowledge distillation, and incur substantial communication overhead. To address these challenges, we propose FedMAGE, a graph hypernetwork-based multimodal federated learning framework. FedMAGE represents client models built from a predefined operator vocabulary as computational graphs and employs a server-side hypernetwork to generate architecture-specific parameters, enabling implicit parameter sharing across heterogeneous clients without requiring a shared parameter space. To improve the generalization of a shared decoder across diverse operator types, we further introduce Type-aware Decoding, which conditions the parameter generation process on operator-specific information. We conduct extensive experiments on three multimodal tasks: image-text retrieval, sentiment recognition, and scene recognition. Compared with a broad range of federated learning baselines, FedMAGE consistently achieves superior performance, improving accuracy by an average of 3.8 % over the strongest baseline, while substantially reducing communication time compared with existing multimodal federated learning methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.