EXIST: Explicit-Guided Spatiotemporal Supervision for EEG Emotion Recognition
Abstract
Emotion recognition in active emotional expression scenarios, such as clinical interviews, relies on audio and video, which provide rich affective cues, but cannot be collected from privacy-sensitive patients. EEG offers a privacy-preserving alternative, yet EEG-based emotion recognition falls behind audiovisual systems due to low signal-to-noise ratio and strong inter-subject variability. We propose explicit-to-implicit supervision, a distillation paradigm that leverages explicit modalities (video and audio) to supervise the implicit modality (EEG) during training and discards them at inference time. Current distillation methods cause negative transfer in this context because they align features or logits across incompatible spaces, and neglect the spatiotemporal structure of the EEG. Therefore, we propose EXIST (EXplicit-guided Implicit SpatioTemporal supervision for EEG), which performs EEG-native reasoning using audiovisual signals as structural supervision within the EEG manifold, enabling EEG to-EEG transfer while avoiding modality conflict. In particular, we design two components: Cortical Dependency Distillation (CDD) and Emotional Dynamics Distillation (EDD). CDD transfers emotion-relevant cortical dependency priors inferred from audiovisual context into EEG-native spatial reasoning, whereas EDD focuses on learning emotionally salient temporal dynamics rather than enforcing direct imitation across modalities. Experiments on EAV and PME4 demonstrate that EXIST achieves state-of-the-art EEG-only emotion recognition under a subject-independent protocol, outperforming both prior EEG-only and existing cross-modal distillation methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.