acceptodds
Under review as a conference paper at ICLR 2027

Speaker-dependent Emotional Transition-aware Multimodal Dynamic Anchor Learning for Multimodal Emotion Recognition in Conversations

Abstract

Multimodal Emotion Recognition in Conversations (MERC) is a crucial task for improving a machine's emotional understanding. It aims to accurately infer emotions from the complex interplay of textual, acoustic, and visual signals in conversational context. Although existing methods have improved performance via advanced strategies like multimodal fusion, graph-based reasoning, etc., they often leave psychologically meaningful multimodal priors and speaker-aware emotion transitions out of consideration to guide multimodal fusion and alignment. In this paper, we propose a novel learning framework, Speaker-dependent Emotional Transition-aware Multimodal Dynamic Anchor Learning (SEDAL), which introduces learnable class-level multimodal semantic anchors and speaker-aware emotional transition dynamics into the multimodal fusion and emotion recognition process. Specifically, SEDAL first constructs momentum-evolving multimodal emotion anchors, which serve as emotion-specific prototypes to guide the multimodal alignment among the text, audio, visual, and fused features. Then, SEDAL further extends these anchors from isolated utterances to adjacent conversational turns through speaker-aware emotion transition prototypes and applies anchor-guided modality perturbation consistency learning only on transition-relevant utterances. Finally, SEDAL further enhances recognition robustness with a deferred student-adaptive multi-teacher boundary distillation strategy, which transfers effective decisions from multiple validation-calibrated teachers. Extensive experiments on IEMOCAP and MELD datasets demonstrate that SEDAL achieves state-of-the-art performance, confirming its effectiveness for MERC.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.