A Cross-Description Test of Sparse Feature Correspondence Between Language Models and Brain Responses
Abstract
Language-model features can predict brain responses. Do the features that matter most for this prediction also matter for predicting the model's own hidden state? We compare sparse-autoencoder (SAE) features using fMRI from 59 people listening to two narrations. We select features and fit separate predictors of fMRI responses and the model's final hidden state on one narration, evaluate on the other, and repeat in reverse. We subtract each feature's contribution from the fixed predictions and compare the signed decreases in held-out variance explained. Across six model–SAE pairs and both directions, mean rank correlation is -.056, and no model passes both timing and feature-identity tests in both directions. Simulations show why this null is inconclusive. At 256 features, known positive correspondence is detected in only 1.6%–10.4% of replicates across 64 conditions from two model families, including planted correlation .80. The false-positive fraction is set to 5%. An exploratory sweep recovers median correspondence of .56 at 24 features and .02 at 256, despite the same planted .80. Separate controls recover the fMRI predictor's planted ranking much better with Gaussian predictors of the same shape. These simulations test reliance estimation with fixed predictor matrices, leaving the full transfer procedure's sensitivity unmeasured. The empirical null cannot rule out shared brain–model organization: prediction and repeatability need a separate check that the analysis can recover known feature agreement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.