Rethinking EEG-to-Image Decoding: From Fine-Grained Temporal Encoding to Hierarchy-Consistent Design
Abstract
Existing EEG-to-image methods commonly use specialized encoders, cross-modal alignment, and generative decoders, but it remains unclear what each stage actually requires or whether added model complexity helps. In this work, we systematically revisit each stage of this common pipeline and introduce TRACE (Temporal Resolution-Aware Complementary EEG Decoding) based on our analysis. At the encoder level, controlled comparisons reveal a counterintuitive trend where a minimal MLP outperforms Transformer-based and large pretrained baselines; our further analysis shows that decoding performance depends strongly on fine-grained temporal encoding, whereas increasing encoder capacity provides no systematic benefit. At the alignment and decoder levels, our analysis identifies unclear separation of level-specific evidence and hierarchical conditioning mismatches. We therefore propose Complementary Dual-Subspace Alignment to organize semantic and structural visual cues into complementary subspaces, alongside Hierarchy-Consistent Injection to route them through a coarse-to-fine visual autoregressive (VAR) generator. On THINGS-EEG, TRACE achieves state-of-the-art retrieval and reconstruction performance. Averaged across ten subjects, it achieves the highest Top-1/Top-5 retrieval accuracy (77.5%/96.2%) and improves six of seven reconstruction metrics over prior methods. Code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.