Selective Coarse-Layer Discrete Distillation for Continual EEG Pretraining
Abstract
Pretrained EEG models may be adapted continually as data arrive from heterogeneous domains. For models that use discrete tokenizers, this adaptation risks both catastrophic forgetting and drift in the token assignments consumed by downstream models. We propose CAP-CL (Coarse Assignment Preservation for Continual Learning), which exploits the hierarchy of residual vector quantization to regularize coarse assignments while allowing finer representations to adapt. CAP-CL distills coarse-layer (L0) assignments from the previous-stage tokenizer on current and replay data, with replay capped separately for each domain and finer layers exempt from the distillation loss. On a stream of 21 EEG datasets, using an autoregressive readout trained with replay on hard token indices, CAP-CL achieves a mean final balanced accuracy of 0.421. It outperforms the tested replay and distillation controls in all five paired seeds, across three fixed stage orders, and under matched coarse-layer loss weights. Preserving assignments between successive stages has different effects from anchoring them to the initial tokenizer: a fixed initial teacher reduces cumulative coarse-layer flips from 0.876 to 0.568 but lowers final accuracy to 0.402. These results support coarse-only discrete distillation as an effective regularizer for continual EEG pretraining with an adaptive readout and show that stronger long-term index preservation alone does not explain downstream performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.