acceptodds
Under review as a conference paper at ICLR 2027

Competence Is Not Control: Causal Decomposition in Multimodal Sentiment Analysis

Abstract

Multimodal sentiment analysis combines video, audio, and text to infer sentiment. Prior work has reported uneven use of these modalities, with models favoring text over visual and acoustic cues. We examine whether correct video-only predictions are preserved when audio and text are added, and how joint input changes video's influence. We focus on cases with correct video-only predictions but incorrect audio-only and text-only predictions. Across three models and three datasets, joint input yields error rates of 51.4%–87.6% on these selected training examples and 83%–97% on neutral examples. Video replacement and gradient analyses show reduced visual sensitivity under joint input, although video still improves accuracy over audio and text alone. On Qwen2.5-Omni, transplanting video-only states read by attention at the answer prompt into runs with all three inputs shifts scores toward the correct label, even with attention weights fixed. Allowing attention to adjust strengthens this effect, with less attention to the answer prompt and more to video, audio, and text on average. With the receiving run fixed, changing the source video has a larger effect for states formed from video alone than under joint input. Matching update size substantially reduces but preserves this difference. For neutral examples, it persists under two update-size controls across two models and two new datasets, for both correct and incorrect joint predictions. On Qwen2.5-Omni and MiniCPM-o 4.5, states from other videos with the same annotated sentiment also shift scores toward the correct label on average. Class-averaged states produce class-specific effects even after matching update size. Together, these findings distinguish a modality's standalone correctness from its influence on joint predictions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.