acceptodds
Under review as a conference paper at ICLR 2027

Extendible Multi-Stage Federated Multi-Modal Learning with Missing Modalities

Abstract

In federated multi-modal learning, communicating the entire model every round increases communication cost and, more importantly, limits how easily the frame- work can be extended, even though handling missing modalities, optimising the downstream task, and, in some settings, training the encoders are each essential parts of the pipeline. Rather than a single jointly trained model, this paper proposes a multi-stage federated learning view in which the component that handles missing modalities is separated from the downstream task: the earliest rounds are devoted to optimising the missing-modality handler alone, with only that component communicated between server and clients, while the remaining rounds are devoted to optimising the downstream task alone, with only that component communicated instead. When pre-trained encoders are not available, this view extends naturally to a third stage in which the initial rounds are devoted to training the encoders themselves. Separately, this paper proposes the missing-modality handler itself, which reconstructs each missing modality from whichever modalities are present. The resulting framework is supported by both theoretical and experimental results. Being task-agnostic, the handler applies directly to regression, classification, and multi-label classification alike, and its F1-micro performance exceeds that of state-of-the-art baselines in the majority of evaluated cases. Code is available at https://github.com/for-reviewers-arz/FArms_wArms/tree/main.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.