acceptodds
Under review as a conference paper at ICLR 2027

StimEM: Stimulus-Conditioned Masked Modeling for Eye-Movement Representation Learning

Abstract

Eye movements provide rich temporal information about human behavior and cognitive function, offering a precise, non-invasive means of assessing cognitive and neurological disorders. However, existing approaches typically rely on handcrafted metrics or task-specific supervised models, whose performance and generalizability are constrained by limited training data and model capacity. Inspired by the success of large language models and foundation models across diverse modalities, we introduce StimEM, an eye-movement foundation model based on Stimulus-Conditioned Eye Movement Masked Modeling. StimEM first learns an ocular tokenizer using finite scalar quantization, jointly optimizing waveform reconstruction and the prediction of informative behavioral features, preserving both local signal structure and trial-level information relevant to cognition. We then pretrain a novel stimulus-conditioned transformer to predict masked tokens. Its structured attention architecture captures dependencies within ocular responses and their relationships with external stimuli. A multi-scale masking strategy further encourages learning of local dynamics and longer-range temporal structure. Pretrained on approximately 500 hours of stimulus-driven eye-tracking recordings spanning multiple cognitive tasks, StimEM achieves state-of-the-art performance on six downstream tasks related to cognitive assessment, demonstrating strong representation learning capabilities and generalizability. Our work first demonstrates the value of self-supervised pretraining for eye-movement data and provides a foundation for further research and broader applications in cognitive science and healthcare.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.