acceptodds
Under review as a conference paper at ICLR 2027

Visual Representations in the Human Medial Temporal Lobe During Naturalistic Movie Viewing

Abstract

Neurons in the human medial temporal lobe (MTL) respond to specific concepts, such as a particular person, across very different images. So far, this selectivity has mostly been measured with isolated images and predefined labels. Representations from vision models provide continuous descriptions of complete visual scenes, enabling MTL activity to be studied without such labels. Using the SUMMER dataset, comprising 2,286 units recorded from 29 intracranially implanted participants while they viewed a full-length commercial movie, we decode frame-wise representations of a wide range of vision models from MTL population activity. We show that MTL activity identifies held-out frames above a temporally matched baseline and locates them coarsely in the film, with performance still increasing at the full population size, whereas low-level image statistics cannot be reliably decoded. Importantly, size and diversity of training data affect decodability. For example, for the same ResNet-50 architecture, decoding performance markedly improves when trained with language supervision on larger datasets. Across all visual feature spaces, a parahippocampal cortex population outperforms size- and firing-rate-matched control populations from the remaining MTL, as well as control populations with twice as many neurons. Our results show that human MTL populations, in particular in parahippocampal cortex, carry information about the visual content of movies that can be used to compare representations from vision models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.