acceptodds
Under review as a conference paper at ICLR 2027

Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs

Abstract

Leveraging the universal representations of pre-trained Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has emerged as a promising paradigm for enhancing the capability of brain foundation models. Constrained by the severe scarcity of visually-evoked electroencephalography (EEG) datasets, existing foundation models predominantly focus on LLMs and compromise by aligning neural signals solely with abstract text—a lossy translation that inevitably discards the fine-grained, perceptual details encoded in brain activity. In this work, we propose Generative Visual Grounding (GVG), a framework that visualizes the invisible. Rather than forcing neural signals into text, we employ an EEG-to-Image generative model as a "visual translator" to hallucinate instance-specific proxy images for non-visual EEG. This strategy provides structured visual contexts that allow MLLMs to apply their visual priors to interpret clinical states. To validate this core hypothesis, we first establish image-only alignment with GVG-X-Omni and GVG-Janus. GVG-X-Omni approaches the average performance of the 1.7B-parameter NeuroLM-XL while tuning 170M parameters over a frozen 7B backbone. Image-only GVG-Janus improves over text-only alignment on all six primary tasks, raising average balanced accuracy from 45.89% to 53.46% and surpassing NeuroLM-XL on four tasks. Building on this result, trimodal EEG–image–text alignment further raises GVG-Janus's average to 60.05% under full adaptation, the highest among the compared multi-task models. Both low-rank adaptation (LoRA) and full adaptation exceed NeuroLM-XL on all six tasks. Predicted native visual tokens also support EEG-based reconstruction through frozen decoders, while clinical analyses identify physiological variation retained in aligned EEG features. These results support visual proxy grounding as an effective complement to textual alignment for universal EEG understanding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.