Which Data, Not How Much: Architecture-Dependent Data Selection for EEG Foundation Models
Abstract
Electroencephalography (EEG) foundation models have achieved strong performance across diverse tasks. Despite differences in architecture and training objectives, most share a common strategy: pretrain on increasingly large data pools to learn better general-purpose representations. This strategy is well established in natural language processing and computer vision, but whether more pretraining data consistently improves EEG transfer remains an open question. We examine 16 candidate datasets, 13 downstream tasks, and 202 configurations across LaBraM, CBraMod, and BIOT, keeping the backbone and pretraining epoch count fixed within each architecture. Adding sources can lower average performance: 121 configurations achieve higher mean test balanced accuracy than full-pool pretraining with fewer training windows. All 39 architecture–task combinations have a smaller, higher-scoring configuration; even adding sources from the target's task family or its own training partition need not help. Similarly sized mixtures yield different benchmark averages and opposing responses on related targets. Capacity experiments show recovery under further expansion at fixed Base capacity; increasing LaBraM to Large narrows the full-pool deficit but leaves a six-source subset ahead by 2.84 percentage points. We compare four practical rankings—Progressive, pairwise dataset valuation (PDV), Size, and Coverage—that identify useful mixtures at different scoring costs, with Size and Coverage sharing their rankings across LaBraM and CBraMod. The findings show that source composition matters alongside data volume and support choosing pretraining datasets using appropriate data-selection methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.