acceptodds
Under review as a conference paper at ICLR 2027

MSSD: Multimodal Fusion as Shared-State Dynamics

Abstract

Multimodal sequence models commonly organize temporal processing around modality specific representations and introduce cross source interaction through separate fusion mechanisms. This leaves recurrent memory either partitioned by source or shared only after source observations have already been combined. We study an alternative organization in which heterogeneous observations remain distinct at the interface but act on one evolving multimodal memory. Multimodal Shared-State Dynamics (MSSD) realizes this view with a single modality free matrix state shared by all sources. Source conditioned write and read factors preserve distinct observation interfaces, while Dual-Axis dynamics jointly evolve latent feature and temporal memory structure within the same recurrent state. The same formulation accommodates different observation times without changing the shared recurrent organization. Across four text audio visual benchmarks, MSSD performs competitively with general and task specific multimodal models, while controlled comparisons and representation analyses support the roles of shared recurrent storage and Dual-Axis dynamics. MSSD thus formulates multimodal fusion as the evolution of one shared recurrent memory rather than as an auxiliary interaction stage around separate source trajectories.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.