acceptodds
Under review as a conference paper at ICLR 2027

vNeuro: Foundation for EEG Vision Models

Abstract

EEG foundation models are typically pretrained from scratch on large corpora of neural recordings. Yet experienced clinical neurophysiologists interpret EEG largely by visual inspection, recognizing spikes, rhythmic bursts, and time–frequency patterns by eye. This raises the question of whether representations learned from natural visual data can be reused for EEG. We introduce vNeuro, a framework that adapts a frozen, video-pretrained V-JEPA 2.1 encoder to EEG through a lightweight trainable interface. vNeuro converts multichannel time–frequency representations into visual tokens that preserve electrode geometry, and trains this interface on unlabeled EEG through latent forecasting with distribution regularization, while the visual backbone remains fixed. During downstream training, only the adapter and classification head are optimized. With a ViT-B backbone, vNeuro achieves balanced accuracies of 81.40% on TUAB abnormality detection and 68.35% on TUEV event classification (68.73% with ViT-G), competitive with EEG foundation models pretrained on tens of thousands of hours of EEG. Ablation and representation analyses show that both the EEG-to-visual token mapping and the regularized self-supervised adaptation contribute to performance. These results suggest that pretrained visual encoders can serve as effective EEG backbones without pretraining a large EEG-specific model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.