PIDFusion: Selective Non-Text Supervision via Partial Information Decomposition for Multimodal Sentiment Analysis
Abstract
Multimodal sentiment analysis (MSA) integrates language, acoustic behavior, and visual expression, yet the dominance of language can cause models to underuse informative non-text cues. Applying additional non-text supervision uniformly is also suboptimal because audio and vision do not contribute unique sentiment evidence in every utterance. To address this issue, we propose PIDFusion, a partial-information-decomposition-guided framework that treats text and the joint audio–visual stream as two information sources for the sentiment target. A neural label-information probe combines source and joint label predictors with label-conditioned Sinkhorn transport to characterize redundant, unique, and synergistic information in continuous representations. The probe is optimized on detached features, separating information measurement from predictive representation learning. We then calibrate a sample-wise audio–visual uniqueness score and use it to selectively allocate an auxiliary text-blind regression objective, thereby strengthening non-text representations where they provide distinct sentiment evidence. In parallel, modality-specific and joint audio–visual regression heads provide direct auxiliary supervision to the corresponding representations. Extensive experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS demonstrate the effectiveness of PIDFusion across English and Chinese MSA benchmarks. Ablation studies further show that PID-guided sample selection is more effective than uniform or matched-random allocation and validate the contributions of the proposed training objectives.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.