acceptodds
Under review as a conference paper at ICLR 2027

HEAR: Hierarchical Evidence-Aware Reasoning via Text Conditioning for Multimodal Sentiment Analysis

Abstract

Multimodal sentiment analysis (MSA) seeks to infer human sentiment from textual, acoustic, and visual signals. Existing approaches often rely on feature fusion or cross-modal interaction, where globally salient information may overshadow subtle yet sentiment-discriminative cues. This highlights a key challenge: perceptual saliency does not necessarily imply sentiment relevance, particularly when modalities convey discrepant information. To address this challenge, we propose HEAR, a Hierarchical Evidence-Aware Reasoning framework that shifts multimodal sentiment modeling from direct feature fusion toward structured evidence reasoning. HEAR first extracts fine-grained candidate evidence from complementary temporal and spectral views of acoustic and visual features while preserving shared and view-specific information. It then conditions the evidence on textual semantics to distinguish text-aligned from text-deviating evidence, explicitly preserving both cross-modal consistency and discrepancy. Rather than directly aggregating evidence into a single global representation, HEAR progressively reasons from local cues through relational interactions to global sentiment abstractions. Hierarchical evidence is incrementally projected into pseudo-tokens and integrated with textual embeddings for sentiment prediction using a frozen large language model. Experiments on CMU-MOSEI, CH-SIMS-V2, MELD, and CHERMA demonstrate consistent improvements over strong fusion-based and LLM-based baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.