When Modalities Disagree: Learning adaptive Task-Aware Multimodal Representations
Abstract
Human-centered multimodal understanding integrates text, audio, and visual cues to infer human sentiment and intent. However, semantic discrepancies, conflicting evidence, and cross-modal redundancy make it challenging to identify task-relevant information and determine appropriate modality contributions during fusion. In particular, predictive disagreement alone cannot determine which evidence should guide the final decision. To address this challenge, we introduce InfoMind, a multimodal representation learning framework that combines shared-private semantic decomposition, asymmetric information-bottleneck regularization, and disagreement-aware latent-state reasoning. Specifically, InfoMind learns compact shared and modality-specific representations and jointly interprets task-level evidence and predictive disagreement through a latent task state, which guides sample-specific fusion of the original multimodal features. Comprehensive experiments across four benchmarks, including CMU-MOSI and CMU-MOSEI for sentiment prediction and MIntRec and MIntRec2.0 for intent recognition, demonstrate the effectiveness and generality of the proposed framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.