Cross-Modal Neural Encoding Is Asymmetric: Evidence From Vision-Language Models and Paired Text-Image Functional MRI
Abstract
Vision-language models (VLMs) have been shown to predict neural responses to visual stimuli better than unimodal models, which has been interpreted as evidence that VLMs and the human brain converge on shared, cross-modal representations. However, existing studies supporting this claim have critical limitations due to the functional magnetic resonance imaging (fMRI) datasets they use, because they often present a single modality, bind modalities inseparably (e.g., a video with simultaneous audio), present mismatched stimuli across modalities, or restrict paired stimuli to isolated concepts. We address these limitations by using a recent phrase- and scene-level text-image fMRI dataset built from paired stimuli: each stimulus exists as both a caption and an image, but any given participant encounters it in only one form. Therefore, cross-modal prediction cannot be attributed to simultaneous presentation, enabling a direct test of whether VLMs and the brain share cross-modal representations. Across multiple VLMs and unimodal baselines, we find a consistent asymmetry: text embeddings reliably predict image-evoked brain activity, whereas image embeddings are largely unable to predict text-evoked brain activity above chance level. We then use a residualization analysis to show that removing the cross-modally shared information from VLM embeddings selectively reduces cross-modal encoding performance to a larger extent than within-modality encoding performance. Mapping the embedding spaces onto each other directly, we find this same asymmetry: text embeddings predict image embeddings more reliably than vice versa. Taken together, these results indicate that VLMs and the brain do share cross-modal information, but not symmetrically: cross-modal information may be more directly accessible from text embeddings than from image embeddings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.