Calibrate Before Disentangling: Robust Multimodal Sentiment Analysis under Missing Observations
Abstract
Robust inference under corrupted observations is crucial for Multimodal Sentiment Analysis (MSA), as real-world textual, acoustic, and visual inputs are often incomplete or unreliable. Existing methods typically recover missing information through direct cross-modal compensation, but this paradigm leaves two key issues insufficiently addressed. First, unreliable observations can induce representation shifts that propagate through subsequent cross-modal interactions. Second, cross-modal evidence can facilitate the recovery of shared semantics but is inherently less suited to restoring modality-specific information. Motivated by these observations, we propose the Calibrated Disentanglement Network (CaDiNet), following a “calibrate-then-disentangle” paradigm. CaDiNet first performs reliability-aware calibration through a dual-granularity residual offset mechanism that leverages missingness information to capture global corruption patterns and instance-specific deviations before cross-modal interaction. It then disentangles each modality representation into shared and private components and refines them using component-specific evidence. Shared semantics are enhanced using reliable cross-modal information, whereas modality-specific information is refined using evidence retrieved from modality-homogeneous historical memory. Extensive experiments on four MSA benchmarks demonstrate that CaDiNet consistently improves robustness across diverse missing conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.