DLink: Distilling Layer-wise and Dominant Knowledge from EEG Foundation Models
Abstract
EEG foundation models (EFMs) provide strong representations for diverse EEG decoding tasks through large-scale pretraining and downstream fine-tuning. Through empirical analysis, we observe that (i) while lightweight models remain competitive, they are highly task-dependent. Retaining high-capacity EFMs as inference backbones incurs substantial computational overhead, making knowledge distillation a natural route for model compression; and (ii) direct distillation from a fixed teacher representation underutilizes EFM knowledge, as task-discriminative information is distributed across intermediate layers rather than concentrated in the final layer. These observations motivate DLink (Distilling Layer-wise and Dominant Knowledge), a spectrally guided distillation framework with input-conditioned layer routing for transferring EFM knowledge into compact students. DLink uses a router to aggregate teacher layers for each input, and aligns magnitude and phase spectra to mitigate compression-induced spectral distortion in learned representations. The routed teacher knowledge is internalized by a project-then-compress student; the teacher and router are used only during training. Experiments across six EEG benchmarks show that DLink improves compact students while remaining lightweight at inference. Further analyses support learned layer aggregation and spectral distillation, establishing DLink as a practical framework for distilling layer-wise dominant knowledge into compact EEG models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.