Alignment of visual brain encoding models with individual perceptual similarity judgments
Abstract
Encoding models that predict brain activation responses to various stimuli, including images, have traditionally been trained to maximize accuracy when predicting the observed brain data. However, in doing so, their internal representations may not reflect individual human perception. For example, two images that have very different predicted brain responses may be perceptually similar to a specific person. In this work, we investigate whether we can fine-tune brain encoding models to better align their representations with specific individuals' similarity judgments. This improved perceptual alignment may be useful in downstream tasks using the brain encoding model, particularly decoding of seen or imagined images. Using measured similarity judgments of 100 images from the 8 participants in the Natural Scenes Dataset (NSD), we adapt a generic human similarity judgment model (DreamSim) separately for each NSD participant to obtain their personalized perceptual similarity models. These personalized perceptual models are then used to obtain subject-specific similarity structure across the images in their encoding model training set. We fine-tune each participant's baseline encoder in two separate conditions, aligning its representations with either the generic or the personalized similarity judgments, with the same initialization and objective. We find that baseline brain encoders reflect within-subject perceptual similarity more strongly than measured fMRI responses and strongly preserve intersubject correlations present in measured fMRI, but fail to reflect the intersubject correlations of human perceptual similarity judgments. Fine-tuning brain encoding models with generic, population-level perceptual similarity information results in even better alignment with human perceptual judgments compared to baseline encoders, but these models still fail to capture intersubject correlation of human perceptual judgments. However, fine-tuning brain encoding models with personalized perceptual similarity information results in models that are overall the most aligned with human perceptual similarity judgments out of all the brain encoders, and also reflect intersubject correlation of human perceptual similarity judgments. Our framework provides a robust approach for bootstrapping human similarity judgment data into creating more individualized, human-aligned brain encoding models of vision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.