Exploring Divergence Between Brain Responses and Language Model Representations Using Crosscoders
Abstract
To what extent do human brains and language models (LMs) share internal representations of language, and how do these representations differ? Prior work has shown that LM representations can predict brain responses to naturalistic language stimuli, suggesting that the two systems encode common information. However, it has remained unclear which features are shared between brain responses and LM representations and which are specific to one of them. We propose Brain-LM crosscoders, which decompose brain responses and LM representations into a common set of features and label each feature as shared, a brain-specific candidate, or an LM-specific candidate based on its predictive contribution to brain responses and to LM representations. Experiments on fMRI data recorded during naturalistic story listening generate hypotheses about information that only brain responses or LM representations encode and that encoding models ignore. Brain-LM crosscoders move the comparison of brains and LMs from prediction to explanation and could help detect concepts that diverge from those of humans.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.