acceptodds
Under review as a conference paper at ICLR 2027

Dissecting the Score: Calibrated Diagnosis of EEG Foundation Models

Abstract

EEG foundation models (EEG-FMs) continue to achieve strong performance on downstream tasks, but task scores alone do not reveal which properties of the signal drive their predictions. In particular, it remains unclear whether predictions depend on physiologically meaningful features of EEG or on simpler properties such as electrode arrangement and signal amplitude. Input perturbations do not fully resolve this question, because their effects can also depend on the adapter, the readout, and artifacts introduced by the perturbation itself. We therefore propose a cross-architecture diagnostic benchmark that uses neurophysiology as a common reference. For each model, we first fix the adapter, pretrained backbone, and readout, and then apply eight diagnostic protocols to the input without further training. Paired controls are used to separate the intended intervention from major side effects, and three representative protocols are validated on synthetic EEG with known signal dependencies. Across five released models and four tasks, we observe a recurring pattern: predictions are often more sensitive to global or low-level signal properties than to the specific brain regions and EEG rhythms expected to be relevant to the task. On motor imagery, swapping corresponding electrodes between the two hemispheres reduces AUROC by 0.18-0.28, whereas completely masking the predefined central sensorimotor region changes performance by less than 0.04. On seizure detection, two models that preserve input amplitude remain highly predictive even after the original waveform is fully replaced by RMS-matched background, yet fall below chance when the original waveform is kept but strongly attenuated. The same design also allows us to test where these behaviors can be changed. With the pretrained backbone kept frozen, a mirror-augmented readout refit makes the motor imagery predictions respond more consistently to left-right electrode swapping. This shows that some behaviors revealed by the diagnostic tests can be changed through the downstream interface without modifying the pretrained backbone. Overall, the benchmark provides a controlled way to uncover signal dependencies that are hidden by task scores and to test whether they can be modified at the downstream stage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.