Concept Generalization in Visual Brain Decoding via Encoder-Generated fMRI
Abstract
Visual brain decoders are typically trained and evaluated on stimuli drawn from similar semantic distributions, leaving their ability to reconstruct unseen concepts poorly understood. We study this concept generalization problem by withholding semantic concepts from a target subject's training data. We find that three recent fMRI-to-image decoders degrade substantially for the held-out concepts, even when the concept remains available for multi-subject models. In contrast, a brain encoder trained to predict fMRI from images is markedly more robust to the same semantic holdout. Motivated by this asymmetry, we use the encoder to generate synthetic fMRI responses for new images and add the resulting image-fMRI pairs to decoder training. Across three decoders, four subjects, and twelve semantic concepts, synthetic data recovers a substantial fraction of the performance lost on unseen concepts without requiring additional, time-consuming fMRI acquisition. We further show that the same approach improves general decoding performance across a wide range of data budgets. Analyses of encoder training conditions and synthetic-data controls show that the gains depend on both encoder generalization and the semantic content of the synthetic supervision, rather than merely increasing the amount of training data. Together, our results show that encoder-generated fMRI can transfer semantic generalization to decoders.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.