From Aggregation to Imagination: Latent World Modeling for Multimodal Federated Learning
Abstract
In multimodal federated learning (MMFL) over nonstationary environments, modality availability couples observation geometry with staleness, producing modality-conditioned temporal misalignment: same-round client evidence may describe different semantic times, and scalar staleness correction cannot remove the resulting directional temporal bias. The remedy is not a better weighting scheme but a different estimand. Federated Latent World Modeling (FedLWM) shifts federation from client-parameter aggregation to a shared latent world state and re-times state inference from server-processing rounds to the semantic times to which the evidence pertains. A fixed hyperspherical semantic reference frame establishes common coordinates, a structured state preserves the geometry on which the dynamics act, and a chart-wise structure-preserving transition supports both retrospective assimilation of delayed evidence and prospective latent imagination at each client's use time. Our theoretical analysis establishes an impossibility result for dynamical prediction from aggregate first- and second-moment summaries: a positive non-closure floor independent of sample size and predictor capacity. It further establishes finite-sample identification guarantees for the structured linear transition blocks and explicit conditions under which retrospective assimilation and prospective imagination improve upon persistence under a shared rollout-error budget. Across four benchmarks, FedLWM consistently outperforms strong baselines under modality-conditioned staleness, with gains widening as modality–staleness coupling intensifies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.