Fusing Evidence, Learning Across Views: Multimodal Sentiment Analysis with Incomplete Data
Abstract
Multimodal sentiment analysis (MSA) is vulnerable to fine-grained missingness, where different corruption patterns retain different subsets of evidence from the same underlying sample. We propose Cross-view Evidence Learning and Alignment (CELA), a same-sample multi-corruption framework that aggregates evidence within each incomplete view and coordinates representations across independently corrupted views. CELA integrates modality-specific and common evidence through slot-based fusion, anchors incomplete-view representations to a task-supervised full-view representation, and explicitly promotes consistency among corrupted-view representations of the same utterance. At inference, CELA requires only a single incomplete observation, without explicitly reconstructing missing content. Experiments on MOSI, MOSEI, and SIMS show that CELA achieves the best missing-rate-averaged performance among the evaluated methods, with clear advantages under low-to-moderate missingness. These advantages also persist under fixed- and mixed-missing-rate training protocols, CIM-oriented nonidealities, and component-level physical ReRAM-based CIM execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.