acceptodds
Under review as a conference paper at ICLR 2027

Does It Really Transfer? Auditing Frozen EEG Foundation Models on Small Ear-EEG Data

Abstract

Pretrained foundation models are often evaluated by freezing them and training a small classifier on a new dataset. When the dataset is small and differs from the pretraining data, a high score can come from choices tuned on the test subjects or from cues unrelated to the task rather than from the representation. We propose a protocol that tells whether such a model really transfers. It records every setup choice, retests each result by repeating the search that produced it on permuted labels and by nested cross-validation, checks whether failed channels or age and sex alone predict the label, and injects known signals to confirm that real effects can be found. We apply it to eight EEG foundation models pretrained on scalp recordings and to Synapse, a wearable ear-EEG dataset of 23 subjects with a passive auditory paradigm and clinical labels, which we release. A standard evaluation finds two models that appear to transfer. The protocol shows that BIOT-18ch's score (0.86 AUROC) comes from the channels that failed quality control and falls to 0.46 without them, and that LaBraM's (0.92) does not rely on failed channels but cannot be confirmed: it holds only in a narrow range of input windows and cannot be separated from sex and age at this sample size. On the same recordings, the protocol confirms a transfer of known size injected into a random group of subjects and confirms nothing without it. We release the protocol, code, data and all results.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.