Hyperbolic Representation Learning for Multimodal-to-Unimodal Transfer in EEG-based Auditory Attention Decoding
Abstract
Auditory attention decoding (AAD) from electroencephalography (EEG) remains challenging across subjects, datasets, and recording systems. While auxiliary physiological signals can enrich EEG representations, their availability at inference time cannot be guaranteed. To solve this problem, we propose a codebook-guided pretraining framework that leverages paired EEG and auxiliary signals to learn transferable representations for EEG-only AAD. During pretraining, separate encoders map the two modalities into hyperbolic space, a hyperbolic residual codebook represents shared physiological structure with common codes and modality-specific information with residual codes. The shared center alignment objective further brings the two modalities' common representations into agreement. The framework then masks parts of the EEG representation and trains the EEG encoder to predict codebook targets learned from the paired multimodal data, thereby transferring the shared vocabulary to the EEG modality. For downstream AAD, the pretrained EEG encoder extracts features from EEG, which are passed to a classifier for attention prediction. Cross-subject and cross-dataset experiments on the AVGCAAD and EARAAD datasets demonstrate that the proposed model outperforms a range of state-of-the-art baselines, and ablation studies confirm that each key module contributes to the performance gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.