No Single Recipe for EEG: Phase, Task Regime, and Data Budget Interact
Abstract
For designers of electroencephalography (EEG) foundation models, input choice depends on the target task and available labeled-data budget. We compare raw patches, Morlet log power, and complex coefficients on P300 and N170 event- related potential tasks and one-second sleep-segment classification across five labeled-subject counts and three seeds. We hold the transformer body, subject splits, epochs, and optimization schedule fixed; only the front end varies. Token counts and compute costs differ. On the ERP composite, raw patches outperform fixed Morlet log power throughout the tested range, with a 0.09–0.12 balanced-accuracy advantage after matching normalization scope. Matching token count and per-example compute in both directions retains the log-power advantage at 20 and 40 subjects; the advantage diminishes at larger counts. Complex coefficients from the same filterbank reduce the evoked deficit. In the ERP composite, the largest decomposition contrast is per-trial, per-frequency phase rotation: removing it improves balanced accuracy by 0.12–0.20. Frame averaging changes ERP-composite accuracy by at most about 0.006 at the same output rate. The rotation preserves pre-normalization magnitudes but changes eventalignment and cross-frequency relationships. A tested convolutional spectral embedding does not remove the evoked deficit. With native-target masked pretraining, raw gains more over scratch than complex inputs, although the arms predict different targets.Under a shared target, the six-seed experiment establishes neither a directional advantage nor equivalence between their gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.