FedM3G: Federated Multimodal Multi-task Graph Learning
Abstract
Federated learning over multimodal attributed graphs lets clients with complementary modalities and heterogeneous downstream tasks collaborate without exposing private data. Under the emerging public-dataset paradigm, clients exchange representations computed on a shared graph rather than model parameters. However, we identify two fundamental obstacles overlooked by current federated multi-task and graph learning methods. First, prevailing multimodal graph encoders entangle node-intrinsic semantics with neighbor-derived topology, so that representation alignment indiscriminately mixes two heterogeneous information streams. Second, task-agnostic alignment forces representations from unrelated tasks onto a single manifold, triggering severe negative transfer. To address these challenges, we propose FedM3G, a federated multimodal multi-task graph learning framework that decouples and selectively aligns representations. On the client side, Gram–Schmidt semantic–topological decoupling exposes the semantics-orthogonal residual of neighbor context and softly combines it with task-relevant aggregated information, yielding a semantic track and a topology-oriented track without imposing rigid orthogonality on their final outputs. On the server side, Task Kernel Alignment estimates label-free public-probe compatibility between clients, and personalized dual-track synergistic representation alignment converts the scores into client-specific teachers and confidence-weighted objectives, with sparse curriculum masking suppressing weak peer relations. Across different datasets, FedM3G consistently outperforms the strongest baseline, with absolute gains reaching up to 16.21% on cross-modal alignment and 4.99% to 9.21% under joint task and domain heterogeneity across five downstream tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.