SAIL: Sparse Autoencoders for Interpretable Alignment of Human Vision and Multimodal Large Language Models
Abstract
Alignment between human high-level visual representations and those of multimodal large language models (MLLMs) offers a quantitative framework for understanding information processing in human vision. However, similarities between model and brain representations do not directly reveal the semantic content underlying their correspondence, as this content is difficult to isolate in entangled MLLM representations. Here, we introduce SAIL, an interpretable framework that formulates voxel interpretation as a multi-unit sparse semantic correspondence problem between MLLMs and human vision. First, we train unified cross-layer sparse autoencoders (SAEs) on MLLM representations of Natural Scenes Dataset images and captions, and align each brain voxel with multiple SAE units. Second, we quantify semantic similarity between their top-response image sets to map semantic alignment across the cortex. We project multiple semantic labels back onto the cortex to interpret the content of this alignment. Finally, we assess these correspondences using category responses measured in independent fLoc experiments. Compared with SAEs trained independently at each layer, SAIL exhibits lower redundancy in SAE unit response profiles and higher semantic alignment. We find that semantic alignment is organized along the human ventral visual pathway across the models and subjects examined. The projected labels reveal semantic preferences within canonical visual regions, while fLoc experiments show agreement between predicted and measured category response patterns. Exploratory analyses further reveal cortical distributions of action-related semantics consistent with established functional accounts of lateral occipitotemporal and posterior parietal cortex. Overall, SAIL provides an interpretable account of the semantic correspondences between MLLM representations and human visual responses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.