acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Vision-Brain Alignment via Layer-Sensitive Temporal Dynamics

Abstract

Continuous neural recordings provide rich temporal dynamics for visual decoding, yet conventional cross-modal alignment methods typically aggregate time-series neural signals into static feature vectors. Such static temporal aggregation neglects the temporal progression inherent in visual cognitive processing, thereby limiting the explicit separation between low-level sensory details and high-level semantics. To address this limitation, we propose Layer-Sensitive Temporal Alignment (LSTA), a progressive framework designed to model continuous temporal dynamics without static feature compression. By structuring neural recordings into a holistic representation alongside early stimulus-evoked and late residual-semantic branches, the framework effectively decouples global context from time-varying perceptual signals. LSTA couples neural response latency with visual hierarchy depth, systematically aligning early perceptual dynamics with shallow visual layers and late cognitive responses with deeper representations. Extensive experiments demonstrate the effectiveness of LSTA, achieving state-of-the-art performance in zero-shot vision-brain retrieval tasks, with Top-1 and Top-5 accuracies reaching 91.4% and 98.6% on the THINGS-EEG dataset, and 42.7% and 67.4% on the THINGS-MEG dataset, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.