When Masked Reconstruction Becomes Spatial Interpolation: What EEG Foundation Models Miss
Abstract
EEG foundation models pretrained by masked reconstruction often trail smaller supervised models, and their gains may depend on task-specific heads or fine-tuning rather than on the pretrained representation. This paper investigates what masked reconstruction transfers to downstream EEG tasks, and why it behaves differently from images. A standard criss-cross Transformer is pretrained on 5,200 hours of EEG data and evaluated on seven benchmarks under linear probing and fine-tuning with a standardised head. When predicting masked EEG patches, a linear map from the visible electrodes at the same time step recovers 66–84% of the encoder's improvement over a simple predictor, so the pretext task can largely be solved by spatial interpolation. Accordingly, during pretraining, spatial attention rapidly converges to a largely trial-invariant map concentrated on neighbouring electrodes, which the same architecture trained on task labels does not form. In contrast, ViT-MAE on images exhibits content-dependent attention under the same analysis. On BCI IV-2a, residual-stream probes show high subject decodability (0.93) but weak class decodability (0.30); across datasets, per-electrode band power is strongly represented before mean pooling but partly lost by it, while channel covariance is substantially less accessible. In a released EEG foundation model, 99% of the pooled representation's energy lies in a component shared across trials, comparable to random initialisation, and falls to 30% after fine-tuning. Replacing learned spatial attention with fixed maps has little effect on downstream performance on most benchmarks, whereas adding covariance and differential entropy features improves performance by up to 11 percentage points on per-subject motor imagery. Together, these results diagnose a modality-specific spatial shortcut in masked EEG pretraining and motivate objectives that explicitly target input-dependent cross-channel structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.