Whose Brain Is It? Reader Identity Accounts for Almost All of EEG-to-Text Retrieval, and a Per-Reader Adapter Recovers It
Abstract
EEG-to-text systems are evaluated on splits that hold out sentences but keep every reader in training, so the reported numbers describe a new sentence read by a known person. We evaluate the setting a deployment would face, a reader absent from training. Tiling all 30 ZuCo readers into five disjoint folds under duration-matched candidate pools, retrieval on a held-out reader is 0.0458, 1.10× chance and not distinguishable from it, against 0.1090, 2.62× chance, for readers present in training on identical pools: 94% of the above-chance margin is reader identity rather than language. The same protocol on ChineseEEG, a different language, paradigm and laboratory, loses 96%, and no label-free normalisation moves it. We then close the gap. Freezing the reader-agnostic encoder and giving each new reader a residual linear remap of the channel axis, at most 16,384 parameters and exactly the identity at initialisation, reaches the seen-reader ceiling on ZuCo from 400 labelled sentences and beats full fine-tuning at matched labels on ZuCo and on clause-level ChineseEEG. Four controls isolate the mechanism: matching the learning rate does not reproduce it, a shared readout with no per-reader parameters is the worst arm, the adapter on a random frozen encoder stays at chance, and a transplanted adapter loses the entire gain. Pre-training and calibration are separately near-worthless and jointly super-additive. Both halves survive a 4 Hz high-pass; the advantage of ZuCo's best electrode subset, which is entirely frontal, does not, a confound we report against our own recommendation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.