acceptodds
Under review as a conference paper at ICLR 2027

BraVista: A Visual Language Framework for Multi-Task EEG Decoding

Abstract

A key challenge in electroencephalography (EEG) decoding is learning representations that generalize across diverse cognitive tasks, subjects, and recording conditions. While EEG provides a non-invasive window into brain activity across a wide range of applications, its datasets are often limited in scale and highly heterogeneous, making multi-task learning difficult. Recent advances in foundation models have improved EEG decoding performance, yet a fundamental open question remains: how to effectively interface neural signals with these models to enable multi-task learning across datasets. To investigate this question, we introduce BraVista, a visual-language framework that encodes multichannel EEG signals as structured images and enables multi-task learning through instruction-conditioned vision-language models (VLMs). Our approach relies on continual post-training of a general-domain VLM, leveraging its visual and linguistic priors to adapt to neural signals without an additional large-scale EEG self-supervised pretraining stage. We evaluate BraVista across four EEG datasets spanning multiple BCI tasks and problem settings, and use this setup to identify the factors that enable effective multi-task EEG learning. First, we show that the right choice of EEG-to-image representation is critical for adapting general-domain VLMs to EEG. Second, we demonstrate that multi-task learning improves decoding performance in low-data regimes, indicating effective knowledge sharing across tasks. Third, a controlled comparison with the same vision encoder and task-specific classifier heads shows that the language model contributes beyond visual feature extraction. Finally, through controlled perturbations of the EEG signal, we observe a gradual performance degradation under increasing noise and task-specific degradation when individual frequency bands are masked, suggesting that the model relies on EEG-relevant information rather than superficial visual patterns. Together, these results establish BraVista as a framework for multi-task EEG decoding via visual representations. Our findings show that visual representations can serve as a viable interface for multi-task EEG learning when paired with appropriate representation and post-training strategies, highlighting the potential of general-domain foundation models as a scalable basis for EEG signal decoding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.