acceptodds
Under review as a conference paper at ICLR 2027

Multimodal Affective Representation Learning: A Technical Evolution from Feature Disentanglement to Semantic Alignment

Abstract

Multimodal affective representation learning aims to learn discriminative, transferable, and robust affective representations from heterogeneous information, including text, speech, visual signals, physiological signals, and eye movements. This review systematically surveys the field from four perspectives: emotion theories, modality-specific information, representation-learning methods, and application-oriented evaluation. Existing approaches are organized into five families: cross-modal alignment and representation disentanglement, cross-modal interaction and dynamic fusion, context and dialogue-structure modeling, self-supervised and robust learning, and large-model-based and prompt-based learning. We further analyze these methods through six affective representation metrics: cross-modal comparability, complementarity, redundancy control, disentanglement of irrelevant factors, robustness to missing modalities, and sensitivity to dynamics, and examine their roles in affect detection and recognition, affective reasoning, affective response and support, and cross-task generalization. Finally, we summarize current evaluation practices and their limitations, with the aim of moving multimodal affective computing beyond single-task performance toward a more systematic assessment of affective representation quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.