acceptodds
Under review as a conference paper at ICLR 2027

Beyond Summaries: Learning to Read Local Evidence for Incomplete Multimodal Sentiment Analysis

Abstract

Incomplete multimodal sentiment analysis predicts sentiment from partially observed language, audio and visual inputs. Existing approaches for incomplete multimodal sentiment analysis predominantly rely on global representations, yet overlook the heterogeneous performance degradation induced by the loss of distinct local regions. Even under identical missing ratios, the absence of decisive local cues may completely flip the predicted sentiment polarity. To mitigate such drawbacks, this paper presents DREAM, a damage-guided framework for discovering crossmodal local evidence. Through contrastive analysis of prediction risks triggered by paired regional erasures, DREAM quantifies the task-level relevance of multi-scale local evidence while preserving global summary information. Equipped with a summary-referenced reader module, each modality adaptively balances a source modality’s global summary with its complementary local evidence, conditioned on the receiving modality’s available content. Extensive evaluations on MOSI, MOSEI and CH-SIMS demonstrate consistent improvements in both sentiment classification and regression. Performance gains are especially substantial on CH-SIMS, where missing fragments exert highly uneven impacts on sentiment polarity. Without full-view teachers or supervision from fully observed counterparts, DREAM improves binary classification accuracy and F1 over MIDAS by 4.11 and 5.05 percentage points, respectively, while reducing MAE by 12.72%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.