acceptodds
Under review as a conference paper at ICLR 2027

Learning Cross-Modal Generative Priors for Multimodal Federated Learning with Missing Modalities

Abstract

Multimodal federated learning often involves clients that observe only one modality, while genuine image–text supervision is concentrated at paired clients. We present FedCGP, a framework that learns conditional feature priors from paired observations and uses them to regularize incomplete clients. Its central idea is to train the representation map actually used during local learning: fitting a denoiser on corrupted targets alone does not ensure useful outputs when queries are instead seeded by independent Gaussian noise. Paired clients therefore combine real-pair alignment, conditional denoising, and contrastive guidance through a single-step query. Incomplete clients keep the prior fixed, detach its outputs, and update their observed encoders alongside real supervised objectives. Block-specific aggregation preserves paired-only prior learning while balancing paired and incomplete contributions to shared representations. We analyze how guidance changes normalized features and how auxiliary gradients interact with real-task learning. Experiments on Flickr30k and MS-COCO combine federation-level evaluation with controlled generator and query-path studies. In the frozen-feature study, generated-to-real supervision substantially improves denoiser correspondence, while deterministic translation remains a strong alternative. Replacing the trained single-step query with a different query path at fixed weights causes a large drop in correspondence, showing that the query procedure is part of the learned interface. These results highlight query-aligned learning and selective update ownership as key design choices for transferring paired supervision under missing modalities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.