acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking Identity-Centric Emotion Tracking in Continuous Video Conversations

Abstract

Current multimodal conversational emotion understanding relies on a convenient but unrealistic premise: pre-segmented utterances with aligned video, audio, and text, where models reason only about the current speaker. In reality, conversations are continuous and unsegmented, and non-speakers also react emotionally. To address this gap, we introduce Identity-Centric Emotion Tracking (ICET), which removes these assumptions. ICET comprises three interdependent sub-tasks: Utterance Segmentation, Identity-Consistent Emotion Classification, and Emotion-Cause Extraction. Utterance Segmentation is performed over all participants, while the latter two sub-tasks cover both speakers and non-speakers. To support ICET, we contribute EmoTrack, a Chinese-English bilingual dataset with 22,585 utterance-level emotion-cause annotations over 21 hours of conversational videos. Benchmarking Multimodal Large Language Models (MLLMs) on ICET reveals challenges in utterance segmentation, long-range dependency modeling, and visual utilization. We further show that speech-aligned transcripts amplify both gains and errors, while reinforcement learning remains unstable due to sparse rewards and credit assignment challenges. These findings highlight EmoTrack's diagnostic value and show that ICET remains an open challenge.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.