acceptodds
Under review as a conference paper at ICLR 2027

When Modalities Disagree: Learning adaptive Task-Aware Multimodal Representations

Abstract

Human-centered multimodal understanding integrates text, audio, and visual cues to infer human sentiment and intent. However, semantic discrepancies, conflicting evidence, and cross-modal redundancy make it challenging to identify task-relevant information and determine appropriate modality contributions during fusion. In particular, predictive disagreement alone cannot determine which evidence should guide the final decision. To address this challenge, we introduce InfoMind, a multimodal representation learning framework that combines shared-private semantic decomposition, asymmetric information-bottleneck regularization, and disagreement-aware latent-state reasoning. Specifically, InfoMind learns compact shared and modality-specific representations and jointly interprets task-level evidence and predictive disagreement through a latent task state, which guides sample-specific fusion of the original multimodal features. Comprehensive experiments across four benchmarks, including CMU-MOSI and CMU-MOSEI for sentiment prediction and MIntRec and MIntRec2.0 for intent recognition, demonstrate the effectiveness and generality of the proposed framework.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.