Visual-to-M/EEG Brain Encoding with Time-Frequency Dual-Path Finite Scalar Quantization
Abstract
Visual brain encoding predicts neural responses from visual stimuli, advancing our understanding of visual perception. It is also an indispensable step in visual prostheses, where predicted neural responses guide electrode stimulation patterns. Prior visual brain encoding research has primarily focused on fMRI, whose limited temporal resolution restricts its applicability. While recent work has shifted toward M/EEG-based brain encoding with diffusion models, these approaches treat M/EEG as images—neglecting their frequency-domain structure—and do not explicitly leverage real neural responses in the training objective. We propose a three-stage visual brain encoding framework. In Stage 1, we propose a novel VAE, called BrainVAE, with a time-frequency dual-path encoder: the input M/EEG signal is treated as a 2D image (channels time) and processed by a Conv2d backbone to extract local spatial-temporal features, while simultaneously transformed by FFT to obtain power spectra. The Conv2d feature tokens and Frequency Prefix Tokens are fused by a Transformer via self-attention, incorporating explicit frequency-domain priors. Finite Scalar Quantization (FSQ) discretizes the latent space without codebook collapse or auxiliary losses. In Stage 2, InfoNCE contrastive learning aligns the quantized neural latent space with I-JEPA image embeddings. In Stage 3, a Transformer-based BrainTokenRegressor predicts FSQ token sequences from I-JEPA image embeddings via cross-attention, with losses computed directly against real neural responses in both token and signal spaces. Experiments on THINGS-EEG2 and THINGS-MEG demonstrate the effectiveness of our approach.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.