Explaining Typical Subject-Invariant EEG Patterns of Emotion through Deep Learning Counterfactuals
Abstract
Counterfactual explanations of electroencephalography (EEG) deep learning models may provide actionable insights for downstream applications such as neuroadaptive brain-computer interfaces, user experimentation, and psychological assessment. However, counterfactual changes in EEG affective state classification are often incompatible with target-class distributions and may produce unrealistic EEG patterns. This complicates the interpretation of EEG-based affect models, particularly when they must generalize across participants. We propose a framework for generating counterfactual explanations under distributional and physiological constraints. We start by introducing a joint variational classifier-decoder architecture that combines affect prediction and EEG feature reconstruction with adversarial regularization to reduce subject-unique information. Then, we propose a counterfactual objective that combines a target-class prediction loss with class-conditional density regularization, proximity penalties in latent and reconstructed feature spaces, and physiological constraints on decoded EEG features. This objective explicitly favors target-class typicality while expressing changes in interpretable EEG frequency-band features. In a strict subject-independent setting, our classifier achieves highly competitive 56.04% valence and 57.50% arousal balanced accuracies on the DREAMER dataset. Typicality-constrained optimization improves counterfactual alignment with the learned target-class distribution relative to an objective without this constraint. All counterfactual EEG patterns generated for held-out subjects had distances to the target class within the 95th percentile of those observed for real target-class samples across subjects, supporting subject-independent generation of counterfactual patterns. These results suggest that typicality optimization alongside prediction change can produce more interpretable EEG affect counterfactuals that generalize across participants.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.