Compact Spatiotemporal Representations via Summary Tokens for Human Intracranial Recordings
Abstract
Intracranial EEG (iEEG) provides high-resolution measurements of human brain activity, but building models that transfer across subjects and tasks is complicated by irregular electrode placement and by downstream tasks that rely on information at different spatial scales. Existing iEEG foundation models typically preserve channel-level structure during pretraining while leaving cross-channel aggregation to downstream adaptation, where labeled data are limited. We introduce a factorized spatiotemporal masked autoencoder that preserves channel-level tokens while learning two complementary summary representations: temporal summary tokens that aggregate information within each channel and spatial summary tokens that aggregate information across channels. During pretraining, an information bottleneck enforces reconstruction through these summary tokens, explicitly encouraging them to capture temporal and population-level structure. For downstream decoding, we use the learned summary representations across a range of iEEG tasks. Our results show that explicitly training the summary representations during pretraining improves downstream transfer across tasks with different spatial requirements. Further, we show that these representations capture task-relevant information and selectively aggregate signals from channels that contribute to downstream decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.