CARE-KAN: Conflict-Aware Evidence Routing for Robust Multimodal Sentiment Analysis
Abstract
Multimodal sentiment analysis typically assumes that language, acoustic, and visual modalities provide mutually supportive, complementary evidence. In real-world affective computing, however, modalities frequently diverge due to nuanced emotional expressions, ambient noise, or sensory degradation. Existing methods struggle with this divergence: forcing cross-modal alignment suppresses subtle modality-unique nuances, whereas directly fusing raw conflict features introduces detrimental noise, triggering severe negative transfer on clean benchmark data. In this paper, we argue that multimodal conflict is an informative epistemic signal regarding sample difficulty and channel reliability, but it is intrinsically unsafe to exploit naively. To resolve this dilemma, we propose CARE-KAN, a conflict-aware evidence routing framework that shifts multimodal sentiment prediction from unconditional feature fusion to calibrated evidence aggregation. CARE-KAN operates through three coordinated stages: (1) Tripartite Evidence Decomposition, which factorizes multimodal features into shared consensus, modality-unique complements, and directional conflict vectors; (2) Reliability-Adaptive Routing, which assesses sample-conditioned modality reliability from the joint interaction of these evidence streams to dynamically modulate their predictive weights; and (3) Zero-Gated Residual Fusion, which incorporates the calibrated evidence via a zero-centered gate, safeguarding clean-set predictive power while selectively utilizing beneficial disagreement cues. The resulting representation is fed into a Kolmogorov–Arnold Network (KAN) prediction head to capture complex non-smooth residual decision surfaces. Evaluations on four benchmarks (CMU-MOSI, CMU-MOSEI, CH-SIMS, and IEMOCAP) show that CARE-KAN eliminates negative transfer on clean data while demonstrating superior resilience under acoustic/visual Gaussian noise and modality dropout. Mechanistic analyses further reveal that conflict magnitude reliably tracks prediction difficulty, and our router dynamically attenuates corrupted channels while preserving salient task-relevant evidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.