EmoGPT: Towards Explainable Multimodal Emotion Recognition with Cross-Modal Chain-of-Thought Supervision
Abstract
Understanding human emotions from multimodal signals requires reasoning over potentially conflicting cues across facial expressions, vocal prosody, and language. However, existing Multimodal Large Language Models (MLLMs) rely primarily on categorical supervision or unstructured descriptions, limiting their ability to perform interpretable cross-modal reasoning, especially in cases of semantic–prosodic inconsistency such as sarcasm. In this work, we introduce structured cross-modal Chain-of-Thought (CoT) supervision for multimodal emotion recognition. We present EmoGPT, a framework that learns to generate emotion predictions grounded in explicit reasoning over visual, acoustic, and textual evidence. To enable this, we construct ECoT-65K, a large-scale Emotional Chain-of-Thought dataset, designed to capture modality-specific evidence and their interactions. Beyond supervision, we propose two inductive mechanisms to improve multimodal reasoning. First, Affective Visual Routing (AVR) introduces a mixture-of-experts projector that adaptively routes visual tokens based on emotion-relevant semantics. Second, Emotional Saliency Bias (ESBias) introduces zero-initialized, layer-specific Q/K/V/O offsets that complement LoRA with emotion-sensitive affine adaptation. Extensive experiments on nine benchmarks under three modality configurations show that EmoGPT achieves the highest mean accuracy in every setting, improving the AVT mean by 4.51 points over the previous state of the art; notably, fine-tuning on ECoT-65K alone already surpasses the previous state of the art without any architectural change. A human audit further shows that 94.3% of the retained reasoning annotations are rated as grounded in multimodal evidence, advancing multimodal emotion recognition toward interpretable and robust cross-modal reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.