A Unified Framework for EEG–Video Emotion Recognition with Brain Anatomy and Temporal Priors
Abstract
Recent studies in video- and EEG-based emotion recognition have shown notable progress. However, multi-modal emotion recognition remains largely unexplored, particularly the integration of physiological signals with video. This integration is crucial, as EEG–video fusion combines observable behavioral cues with internal neural dynamics and enables a more comprehensive and robust characterization of human emotion. To this end, we propose EVER, a novel EEG–Video Emotion Recognition framework that effectively integrates complementary information from both modalities. Specifically, EVER models the anatomical organization of multi-channel EEG signals and the temporal relationships among video clips through graph-based representations. These structured features are subsequently integrated with global modality representation using an inter-modal graph for emotion prediction. To provide a comprehensive evaluation, we establish a benchmark across two public EEG-video paired datasets, Emognition and MDMER. We evaluate 12 representative models, consisting of 5 EEG-only, 5 video-only, and 2 EEG-video models. Extensive experiments demonstrate that the proposed EVER achieves state-of-the-art performance by jointly modeling behavioral cues from video and physiological responses, thereby enabling the recognition of emotional patterns unattainable by either modality alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.