Modality-agnostic brain-tuning of text and speech representations
Abstract
The human brain has largely shared higher-level language representations during reading and listening, despite differences in early sensory processing. Recent "brain-tuning" studies have shown that fine-tuning language models with brain data can improve their semantic representations, but these studies have focused on a single modality. Here we ask whether brain-tuning can help multimodal models learn brain-relevant representations that are shared across modalities. We brain-tune two multimodal text-speech models using fMRI data collected during naturalistic reading and listening. We then evaluate how well the brain-tuned model encodes brain activity, both within and across modalities. We further evaluate the brain-tuned model's representation on several semantic and audio probing tasks. All analyses are conducted across model layers, on held-out participants. Our results show that brain-tuning improves both within- and cross-modality prediction accuracy across language-selective brain regions. These gains emerge primarily in middle and late layers, where brain-tuning also strengthens semantic and audio information. Notably, tuning with both modalities jointly yields greater improvements, for both encoding performance and probing tasks, than tuning with either modality alone. Together, these findings suggest that brain-tuning can strengthen brain-relevant language information shared across modalities, providing a way to ground multimodal language models in human brain representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.