acceptodds
Under review as a conference paper at ICLR 2027

From Cues to Segments: Learning Latent Affective Units for Multimodal Temporal Emotion Localization

Abstract

Multimodal temporal sentiment localization aims to identify and localize emotion-bearing segments in untrimmed multimodal sequences. However, cross-modal temporal misalignment and sparse affective cues are challenging for existing fusion methods, which often introduce redundant or unreliable information. To address these challenges, we propose LLAU, a multimodal fusion method that learns temporally structured latent affective units to construct dense and boundary-sensitive representations for emotion localization. Specifically, LLAU learns shared and modality-specific units to capture complementary affective cues. By jointly considering cross-modal support, semantic affinity and temporal affinity, it iteratively refines soft token-to-unit assignments and unit representations, thereby organizing sparse multimodal cues into compact affective representations. The learned assignments serve as affective priors for reconstructing dense features along the original timeline and guide modality-specific experts in weighting residual features to mitigate the loss of emotion boundary details during aggregation and reconstruction. We further introduce EmoTEL-8, a target-centric temporal emotion localization benchmark integrating visual, acoustic, and textual modalities. It provides identity-consistent visual tracks of target individuals, complete acoustic and textual context, and precise temporal boundaries for eight emotion categories. Experimental results demonstrate that LLAU achieves state-of-the-art performance on EmoTEL-8 and the public audio-visual benchmark TSL-300.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.