GLANCE: Global–Local Alignment of Neural and Contextual Embeddings for Brain–Language Sentence Retrieval
Abstract
Identifying heard language from neural activity requires models that align distributed, time-varying brain signals with linguistic representations. We introduce GLANCE (Global–Local Alignment of Neural and Contextual Embeddings), a framework for sentence retrieval from stereo-electroencephalography (sEEG) recorded during naturalistic movie viewing. GLANCE models within-lead temporal dynamics using recurrent encoders, combines information across leads with spatial attention, and integrates temporal context into a global sentence representation. A complementary local pathway compares time-resolved neural features with contextual word embeddings from each candidate transcript. Both pathways operate in embedding space, combining whole-sentence similarity with word- level evidence without requiring word-onset supervision or a fixed word-to-time alignment. GLANCE is trained with a multi-positive contrastive objective that accounts for repeated transcripts, lightweight adaptation of frozen text features, and paired full and partially masked neural views. We evaluate the framework on 13 recordings from four Brain Treebank participants, with separate models trained and tested within each participant–movie recording. Under 100-way retrieval, GLANCE achieves 26.2% subject-averaged Recall@10 and exceeds the 10% chance level in every recording. Combining global and local scores generally improves retrieval over global-only scoring across recordings. These findings demonstrate the value of joint sentence-level and word-level matching for retrieving heard sentences from naturalistic intracranial activity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.