Robust Multimodal Sentiment Analysis via Calibrated Dynamic Fusion
Abstract
Multimodal Sentiment Analysis (MSA), which integrates linguistic, acoustic, and visual modalities, is crucial for understanding human emotions. Since modality quality often varies in real-world scenarios, dynamic fusion methods have been introduced to adaptively weight modalities according to their estimated reliability. However, existing methods still suffer from two key limitations: they often confuse degradation caused by intrinsic noise with that caused by insufficient optimization, inducing erroneous modality modulation. Besides, they usually rely on simple probabilistic assumptions that cannot capture the complex, multi-peak structure of real modality features. To address these issues, we propose a Calibrated Dynamic Fusion (CDF) framework for robust MSA. Specifically, we introduce Degradation Source Rectification (DSR) to reduce optimization-induced uncertainty bias, preventing informative but under-optimized modalities from being unduly downweighted. We further propose Multi-Peak Uncertainty Modeling (MPUM) to capture heterogeneous local patterns in modality representations and provide more faithful uncertainty estimates under complex feature distributions. The calibrated uncertainty estimates are then used to assess modality reliability and guide sample-adaptive dynamic fusion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.