STREAM: Sample-Adaptive Two-level Routing Experts with Aligned Mamba for Multimodal Representation Learning
Abstract
Multimodal Sentiment Analysis (MSA) aims to understand human emotions by integrating acoustic, visual, and textual signals. However, existing approaches often rely on static fusion architectures with limited interpretability, overlooking the dynamic and hierarchical nature of cross-modal interactions. To address these limitations, we propose STREAM, an interpretable dynamic fusion framework for MSA. STREAM introduces an input-conditional operator routing mechanism combined with a Partial Information Decomposition (PID)-inspired Mixture-of-Experts design, enabling sample-specific modeling of unique, redundant, and synergistic cross-modal patterns. Furthermore, STREAM incorporates hierarchical representation alignment and a Collaborative Bi-Mamba module to effectively capture interactions across different semantic levels and temporal dependencies. Beyond prediction, STREAM provides interpretable insights by analyzing routing behaviors, expert activation patterns, and modality-specific contributions. We evaluate STREAM on three benchmark datasets, including CMU-MOSI, CMUMOSEI, and CH-SIMSv2. Experimental results demonstrate competitive performance across these benchmarks, achieving state-of-the-art results on CH-SIMSv2 and strong results on CMU-MOSEI, with an Acc-7 of 54.58%, Acc-2 of 85.74%, MAE of 0.531, and Corr of 0.772. By offering both adaptive fusion and interpretable decision analysis, STREAM provides a promising foundation for future affective computing applications, including intelligent human-computer interaction and emotion-aware systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.