Predictive Structure Is Distributed Across Depth in EEG Foundation Models
Abstract
Intermediate layers often outperform final layers when foundation models are used as frozen feature extractors, but the source of this advantage is unclear. We study this question in EEG foundation models by separating three factors: how representations are read out, whether intermediate layers retain predictive features that complement the final layer, and whether useful depth structure transfers across datasets. Across five models and three tasks, improving the final-layer readout consistently recovers substantial predictive performance, yet intermediate layers remain advantageous in most settings. Residual analyses further show that intermediate representations contain predictive variation that complements final-layer features, including cases where the intermediate layer is weaker on its own. This pattern persists with a nonlinear cross-layer mapping. We also find that layers selected on one dataset remain useful on another, and that a compact set of source-derived task directions retains predictive value after transfer. Together, these results show that intermediate-layer utility cannot be reduced to either poor final-layer pooling or better standalone representations. Instead, predictive structure is distributed across depth, and its usefulness depends on how representations are accessed and combined.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.