Encoding Without Alignment: Analyzing The Correspondence Between LLMs And The Visual Cortex
Abstract
Large language model (LLM) embeddings learned only from text can predict human visual-cortical responses to natural scenes, but whether this predictive correspondence implies shared representational geometry remains unclear. Using fMRI responses from all eight Natural Scenes Dataset subjects and Llama-3-8B representations of the corresponding COCO captions, we compare two notions of brain-model correspondence. First, voxel-wise ridge encoding yields positive held-out prediction accuracy across cortical regions, with higher accuracy in body-selective cortex than in early visual cortex and little variation across Llama layers. Second, we characterize local representational geometry within semantic neighborhoods. Principal-component variation patterns over the shared neighboring stimuli define subspaces in a common neighbor-index space, whose orientations we compare using chordal distance. Within each system, tangent-space distance increases with separation along the cortical ROI ordering and with Llama layer separation. Across systems, however, alignment remains close to an isotropic random-subspace baseline, while being modestly stronger in EBA and FBA than in early visual cortex. Local variance spectra also dissociate: cortical effective dimensionality varies comparatively little across regions, while Llama effective dimensionality increases with layer depth. An exploratory residual-stream intervention using encoding-derived directions produces mixed and nonspecific effects across four subjects. In contrast to Llama, CLIP text-encoder representations show increasing cortical encoding accuracy and stronger cross-system alignment with depth. These results demonstrate that accurate linear encoding can coexist with weak alignment of dominant local variation subspaces.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.