Channel-Preserving Representation Alignment for EEG-to-Music Reconstruction
Abstract
Reconstructing naturalistic music from EEG requires extracting music-specific information from weak signals distributed across electrodes. Existing spectrogram regression and diffusion approaches provide audio synthesis mechanisms, but learning the EEG-to-music mapping faces an information–estimation tradeoff. Combining electrode measurements into fewer features can ease estimation from limited recordings while obscuring distinctions between music segments. We propose channel-preserving representation alignment, separating electrode information retention from prediction regularization. Per-electrode tokenization preserves spatial identity for learned aggregation, while multi-view self-distillation and channel dropout encourage consistency across electrode subsets during pretraining and music alignment. The aligned representation conditions a pretrained audio generator. Our theory explains these complementary choices. A finite-sample regression analysis identifies when retained task information outweighs estimation cost. For nonlinear alignment, an exact decomposition reveals a prediction-consistency penalty in the masked contrastive objective, with margin-dependent bounds connecting prediction variability to ranking errors across electrode subsets. On NMED-T/H, our method improves within-subject 50-way identification from 0.388 to 0.485 and generated-audio CLAP similarity from 0.635 to 0.683 compared with previous work.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.