acceptodds
Under review as a conference paper at ICLR 2027

Discern, Then Combine: Probe-Guided Multi-Dataset Fine-Tuning for EEG Foundation Models

Abstract

Electroencephalography (EEG) foundation models pretrained on large-scale heterogeneous corpora provide general-purpose representations, yet the standard practice of fine-tuning them on a single target dataset, while direct and widely used, has proven suboptimal in many settings. While incorporating auxiliary datasets can provide effective multi-domain regularization, indiscriminate data mixing frequently induces negative transfer due to incompatible characteristics. Crucially, dataset compatibility in EEG is highly counter-intuitive and cannot be reliably predicted from surface-level task labels or metadata. We propose POSE, a data-centric, target-conditioned framework for multi-dataset EEG fine-tuning. POSE uses early target-only adaptation as a dynamic probe, combining endpoint discrepancies and temporal trends to select and weight auxiliary datasets. A phased training schedule then regulates auxiliary participation to balance joint learning with target-focused refinement. Comprehensive experiments show that POSE achieves higher average performance than all evaluated baselines and improves over target-only fine-tuning on most targets, while maintaining stable and robust performance across heterogeneous settings. Beyond performance improvements, our selection analyses and case studies provide empirical guidance for deciding when and how to incorporate auxiliary data for a given target and backbone. More broadly, this work advances a compatibility-aware perspective on EEG foundation-model adaptation, encouraging further exploration of how heterogeneous datasets can complement one another in target-specific learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.