Where, When, and What Aligns: A Retrieval-Grounded Atlas of EEG–ViT Alignment
Abstract
The alignment between human brain activity and pretrained vision models has become both a window into biological vision and the engine of EEG-to-image decoding. In this work, we draw the first systematic atlas of the EEG–ViT interface measured in the deployment metric itself, end-to-end 200-way retrieval, resolved jointly over scalp position, epoch time, network depth, sub-layer computation, token readout, and training objective. Analyzing 21 vision models spanning contrastive, supervised, self-distillation, and reconstruction training, we evaluate approximately 21,000 probe cells on all ten THINGS-EEG2 subjects at three trial-count regimes ( averaged test repetitions), with every claim supported by a per-subject sign test. Our analysis reveals three key findings: (1) where and when: the network's depth axis carries a brain timestamp; each block aligns best with a distinct, progressively later moment of the EEG epoch (177 to 201 ms across blocks 0 to 11, in 10/10 subjects), co-registered with the scalp's posterior-to-anterior latency gradient; (2) where in the network: the well-documented mid-depth alignment peak has a mechanism; early blocks write EEG-visible content into the residual stream, the mid-network stream integrates those writes and overtakes every computation, and late computations fade while the stream coasts on what was already written; (3) what: what EEG shares with a vision transformer is low-level visual structure at coarse spatial grain, not fine-grained semantics: alignment is carried by electrodes over early visual cortex, peaks at one-third to one-half of network depth, deeper still under masked autoencoding, and fades exactly where representations turn semantic, in the late layers, unless the training target itself is pixel reconstruction; spatially, a coarse to pooling of the token grid outperforms the field's global readouts by 6 to 22 points and any single patch by far more, the grain expected of a 63-electrode montage reading through the scalp. Building on these insights, the atlas configures decoders by measurement alone: it distills a five-model fusion pipeline to eight electrodes and a single backbone that surpass the published state of the art, and composes a four-state fusion chain that reaches 97.5% final-epoch accuracy together with the first single-trial accuracy above 22%. Our results suggest that the brain–transformer interface is a lawful and largely unexplored territory, and that deployment-currency measurement offers a new instrument for studying biological vision through transformer computations and vision models through the organization of the brain.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.