MoRSE: Distilling EEG Foundation Models into a Single Montage-Robust Multi-Task Encoder
Abstract
Electroencephalography (EEG) foundation models are pre-trained across many electrode montages, and several accept any number of channels by construction. However, they are typically deployed after fine-tuning on a task recorded with a single montage, which largely removes this montage flexibility. Existing remedies, such as channel interpolation or montage-sampled fine-tuning, still leave one large model per task. To address this, we propose MoRSE (Montage-Robust Student for EEG), which distills fine-tuned EEG foundation models into a single small multi-task encoder that remains accurate on electrode subsets. In this asymmetric distillation, teachers observe the full montage while the student observes a random electrode subset resampled each step, and so learns to infer the full-montage representation from the electrodes available. Having found that fine-tuning typically compresses representations to an effective dimensionality that tracks the class count, we align the student with a never-fine-tuned backbone. We evaluate MoRSE on eight tasks spanning seizure detection, sleep staging, emotion recognition, motor imagery and dementia diagnosis, where a model fine-tuned without montage variation loses 0.142 balanced accuracy when electrodes are removed. Across 30 task-by-montage cells, a single 1.97M-parameter encoder retains 95% of the balanced accuracy of the per-cell specialists, with 89× fewer deployed parameters. Against eight montage-sampled specialists that are 24× larger in aggregate, it retains 98%. It also transfers zero-shot to unseen electrode layouts and to two external sleep-staging datasets. Thus, one encoder a third the size of a single specialist can replace multiple task-specific foundation models across tasks and layouts. Anonymous code is available via the Reproducibility Statement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.