MSSD: Multimodal Fusion as Shared-State Dynamics
Abstract
Multimodal sequence models commonly organize temporal processing around modality specific representations and introduce cross source interaction through separate fusion mechanisms. This leaves recurrent memory either partitioned by source or shared only after source observations have already been combined. We study an alternative organization in which heterogeneous observations remain distinct at the interface but act on one evolving multimodal memory. Multimodal Shared-State Dynamics (MSSD) realizes this view with a single modality free matrix state shared by all sources. Source conditioned write and read factors preserve distinct observation interfaces, while Dual-Axis dynamics jointly evolve latent feature and temporal memory structure within the same recurrent state. The same formulation accommodates different observation times without changing the shared recurrent organization. Across four text audio visual benchmarks, MSSD performs competitively with general and task specific multimodal models, while controlled comparisons and representation analyses support the roles of shared recurrent storage and Dual-Axis dynamics. MSSD thus formulates multimodal fusion as the evolution of one shared recurrent memory rather than as an auxiliary interaction stage around separate source trajectories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.