acceptodds
Under review as a conference paper at ICLR 2027

Brain-VisA: Brain-Vision Alignment via Self-Supervised Brain Representation

Abstract

Human brain activity contains rich perceptual and semantic information that could complement pretrained models, yet exploiting this information is challenging. We investigate whether brain–vision alignment can improve the representations and generalization of pretrained vision models using paired image–fMRI data alone, without any behavioral supervision or category information. To enable vision models to benefit from brain–vision alignment, we leverage self-supervised learning to enhance the conceptual and semantic structure of brain representations. Specifically, we propose a two-stage **Brain-Vis**ion **A**lignment (**Brain-VisA**) framework. First, a joint-embedding predictive architecture (JEPA) learns structured brain representations through self-supervision, using repeated fMRI responses to the same image as natural positive views. Second, relational alignment jointly adapts brain and visual representations while retaining brain self-supervision. We find that the learned brain representations exhibit semantic concept structure across eight pretrained vision models from SimCLR, CLIP, and DINOv2, despite the absence of semantic supervision. Specifically, Brain-VisA improves one-shot classification accuracy for high-level semantic concepts by an average of 13.70% over the unadapted models. It further exceeds the category-informed brain–vision alignment baseline by an average of 10.14% on high-level abstract semantic classification, while requiring less training data and time. The aligned models also improve on generalization tasks including out-of-distribution recognition and the odd-one-out task. These results suggest that self-supervised learning can extract useful representational structure from limited, noisy brain signals, and that brain–vision alignment can improve the semantic representation structure and generalization of pretrained vision models without relying on any behavioral or category information.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.