acceptodds
Under review as a conference paper at ICLR 2027

Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli

Abstract

Multimodal large language models (MLLMs) align better with brain activity than unimodal models. Instruction-tuned (IT) MLLMs, in which a language instruction governs how visual and auditory input is processed, produce task-specific representations. Yet their brain alignment has been evaluated only under unimodal stimuli, while studies using multimodal stimuli have relied on non-instruction-tuned models. It therefore remains unclear whether, under naturalistic multimodal input, instruction-conditioned representations capture task-related structure beyond the surface semantics of the instruction text. We address this by predicting fMRI responses recorded during naturalistic movie watching (video with audio) from MLLM representations. Across three matched pairs of pretrained and IT MLLMs given identical instructions (the pretrained models prompted via in-context learning (ICL)), IT models align better with brain activity at every layer. In all three pairs, ICL representations track instruction-text semantics (–), whereas IT representations are largely decoupled from wording (–) and instead exhibit significant task-conditioned organization throughout the network. This does not mean IT models ignore the instruction: they respond to which task is asked, not to how it is worded. Surprisingly, for identical movie stimuli, different instructions yield distinct regional patterns of brain predictivity. Narrative understanding in the angular gyrus and posterior cingulate cortex, spatial understanding in parietal and scene-selective regions, and emotion understanding in the middle frontal gyrus. These preferences are broadly consistent across participants. Eight video and two audio IT-MLLMs replicate this instruction-dependent regional structure. Together, these findings show that instructions with IT-MLLMs offer a controlled way (model organisms) to derive many task-specific representations from a single stimulus, revealing which cortical regions are best predicted by which task. Because the neural data are held fixed, this provides a low-cost, model-side source of hypotheses to complement task manipulation in neuroimaging, which requires separate acquisitions per task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.