A Patchwise Aggregation Encoder for EEG-to-Image Retrieval: An Analysis of What Image Features Matter
Abstract
What image features matter in EEG-to-image retrieval? Published approaches on the Things-EEG benchmark align EEG responses with features of a pre-trained visual encoder, but which image features drive this alignment remains unclear. We answer this question with the Patchwise Aggregation Encoder (PAE), a 1.57 M-parameter image encoder trained from scratch. The image is split into non-overlapping patches, each patch is processed independently by a shared stack of residual blocks, and the patch features are fused by a learned weighted average. On zero-shot within-subject retrieval over 200 unseen images, PAE reaches a mean top-1 of 96.9% and top-5 of 99.7% across the 10 subjects, a new state of the art on this benchmark. A purely statistical variant, S-PAE, replaces the pixel input with 15 low-level statistics of each patch and is even stronger, at 97.6% top-1 and 99.9% top-5. Ablations then identify the features this alignment relies on. The feature space is dominated by luminance contrast and edge statistics, and the learned pooling weights concentrate on the image center, consistent with foveal dominance in early visual processing. Retraining with only three statistics retains 95.2% top-1 accuracy, and brightness and contrast alone reach 90.3%. EEG-to-image retrieval thus rests on the low-level statistics of local patches, computed independently and combined by learned spatial weights, which makes decoding compact and interpretable by construction. These findings offer a new perspective on EEG-image alignment and motivate further research on EEG-based visual decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.