EmoTrust: Learning Which Modality to Trust for Multimodal Emotion Reasoning under Cross-Modal Conflict
Abstract
Multimodal emotion reasoning aims to infer human affect from visual, acoustic, and linguistic signals, but remains challenging when modalities convey conflicting emotional evidence. Existing methods improve perception and reasoning quality, yet they typically fuse multimodal cues without explicitly modeling which modality should be trusted more when cross-modal emotional conflict arises. To address this challenge, we propose EmoTrust, a framework for multimodal emotion reasoning that explicitly resolves such conflicts through adaptive modality preference learning. EmoTrust begins with multi-view supervised fine-tuning on full and modality-ablated inputs, enabling robust reasoning under varying modality availability. It then estimates cross-modal consistency through cue-grounded modality-level emotion support distributions. To resolve modality disagreement, EmoTrust further introduces conflict-conditioned policy optimization, which compares confidence under full and ablated inputs to estimate modality-specific contributions and reinforce reasoning grounded in the more informative modality. Experiments on several benchmarks show that EmoTrust consistently outperforms strong baselines, especially on low-consistency samples, while learning modality preferences that align well with human judgments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.