acceptodds
Under review as a conference paper at ICLR 2027

Explaining Typical Subject-Invariant EEG Patterns of Emotion through Deep Learning Counterfactuals

Abstract

Counterfactual explanations of electroencephalography (EEG) deep learning models may provide actionable insights for downstream applications such as neuroadaptive brain-computer interfaces, user experimentation, and psychological assessment. However, counterfactual changes in EEG affective state classification are often incompatible with target-class distributions and may produce unrealistic EEG patterns. This complicates the interpretation of EEG-based affect models, particularly when they must generalize across participants. We propose a framework for generating counterfactual explanations under distributional and physiological constraints. We start by introducing a joint variational classifier-decoder architecture that combines affect prediction and EEG feature reconstruction with adversarial regularization to reduce subject-unique information. Then, we propose a counterfactual objective that combines a target-class prediction loss with class-conditional density regularization, proximity penalties in latent and reconstructed feature spaces, and physiological constraints on decoded EEG features. This objective explicitly favors target-class typicality while expressing changes in interpretable EEG frequency-band features. In a strict subject-independent setting, our classifier achieves highly competitive 56.04% valence and 57.50% arousal balanced accuracies on the DREAMER dataset. Typicality-constrained optimization improves counterfactual alignment with the learned target-class distribution relative to an objective without this constraint. All counterfactual EEG patterns generated for held-out subjects had distances to the target class within the 95th percentile of those observed for real target-class samples across subjects, supporting subject-independent generation of counterfactual patterns. These results suggest that typicality optimization alongside prediction change can produce more interpretable EEG affect counterfactuals that generalize across participants.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.