acceptodds
Under review as a conference paper at ICLR 2027

EmoGPT: Towards Explainable Multimodal Emotion Recognition with Cross-Modal Chain-of-Thought Supervision

Abstract

Understanding human emotions from multimodal signals requires reasoning over potentially conflicting cues across facial expressions, vocal prosody, and language. However, existing Multimodal Large Language Models (MLLMs) rely primarily on categorical supervision or unstructured descriptions, limiting their ability to perform interpretable cross-modal reasoning, especially in cases of semantic–prosodic inconsistency such as sarcasm. In this work, we introduce structured cross-modal Chain-of-Thought (CoT) supervision for multimodal emotion recognition. We present EmoGPT, a framework that learns to generate emotion predictions grounded in explicit reasoning over visual, acoustic, and textual evidence. To enable this, we construct ECoT-65K, a large-scale Emotional Chain-of-Thought dataset, designed to capture modality-specific evidence and their interactions. Beyond supervision, we propose two inductive mechanisms to improve multimodal reasoning. First, Affective Visual Routing (AVR) introduces a mixture-of-experts projector that adaptively routes visual tokens based on emotion-relevant semantics. Second, Emotional Saliency Bias (ESBias) introduces zero-initialized, layer-specific Q/K/V/O offsets that complement LoRA with emotion-sensitive affine adaptation. Extensive experiments on nine benchmarks under three modality configurations show that EmoGPT achieves the highest mean accuracy in every setting, improving the AVT mean by 4.51 points over the previous state of the art; notably, fine-tuning on ECoT-65K alone already surpasses the previous state of the art without any architectural change. A human audit further shows that 94.3% of the retained reasoning annotations are rated as grounded in multimodal evidence, advancing multimodal emotion recognition toward interpretable and robust cross-modal reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.