Reason, Decompose, and Deliberate: Partial-Information-Guided Multimodal Sentiment Analysis
Abstract
Multimodal Sentiment Analysis (MSA) is challenging because textual, acoustic, and visual cues relate to sentiment differently across instances. They may jointly convey a meaning absent from any single modality (e.g., sarcasm), or leave the decisive cue to one modality. To handle this variability, we propose the Reason–Decompose–Deliberate (RDD) framework, which organizes multimodal evidence into three cross-modal roles of redundancy, synergy, and uniqueness. This role-level view turns fusion into an instance-specific decision about how evidence should be combined, so that shared, interaction-dependent, and modality-specific cues are modeled separately rather than entangled in a single representation. To Reason, Multimodal Information-Role Reasoning and Routing (MIRR) prompts an MLLM with chain-of-thought to estimate an instance-specific cross-modal role vector and maps it to routing gates. To Decompose, Information-Role Guided Expert Decomposition (IGED) extracts redundant, synergistic, and unique features with dedicated experts and fuses them via gate-weighted summation across modalities. To Deliberate, Uncertainty-Aware Latent Deliberation (UALD) identifies uncertain samples by predictive entropy and iteratively refines their representations with a latent think token before prediction. Consequently, RDD unifies role reasoning, role-specific decomposition, and uncertainty-driven refinement into a single process that adapts to the evidence configuration of each instance. Experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS show that RDD achieves the best results on all reported metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.