Subject Identity as a Shortcut: Rethinking Self-Supervised Pretraining for EEG Foundation Models
Abstract
Electroencephalography foundation models (EEG-FMs) promise to learn transferable neural representations from unlabeled corpora, yet whether self-supervised pretraining achieves this remains unclear. We audit frozen representations of five released EEG-FMs across five canonical brain–computer interface (BCI) benchmarks using linear probing, variance decomposition, -SNE, and representational similarity analysis. A consistent pattern emerges: pretraining strengthens subject identity information without comparable gains in task-related information, revealing as a shared failure mode of prevailing EEG self-supervised paradigms. To identify the underlying mechanisms, we pretrain MAE-style and MoCo-style EEG-FMs on a multi-source corpus. Masked reconstruction preferentially exploits subject-related spatial structure over temporal context, whereas contrastive learning leverages subject differences for positive-pair consistency and separation from predominantly cross-subject negatives. Euclidean-aligned targets fail to suppress this shortcut in MAE, whereas same-subject negatives reduce it in MoCo without reliable task gains. These results motivate EEG-specific objectives that explicitly learn transferable task-related representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.