Temporal Structure in Contrastive EEG Representation Learning: A Controlled Study for Obstructive Sleep Apnea Detection
Abstract
Contrastive self-supervised learning for physiological signals typically treats windows as independent apart from augmentation-defined positives, ignoring that windows from the same patient are temporally related. We introduce Temporal NT-Xent, which retains augmentation-defined positives and full-strength cross-patient negatives while continuously down-weighting the repulsive contribution of same-patient negatives according to their temporal distance. We evaluate this objective for EEG-based obstructive sleep apnea (OSA) detection along three complementary axes. First, on a canonical patient-held-out split—with patients unseen during pretraining, validation, or test selection—Temporal NT-Xent produces a substantial downstream change relative to vanilla NT-Xent, raising AUROC from 0.5417 to 0.7667 and AUPRC from 0.4013 to 0.6673 (paired permutation p = 0.0455 for AUROC). Second, the same frozen 131K-parameter encoder transfers zero-shot to the independent SHHS-1 cohort without any apnea labels or target-domain fine-tuning, reaching AUROC = 0.724 (95% CI [0.641, 0.801]), compared with 0.718 for a single-channel source-to-target control. This probes cross-cohort transfer separately from the in-domain question addressed by the canonical split. Third, a five-seed robustness analysis on the source cohort shows that the magnitude of the in-domain downstream gain is not seed-robust (mean ΔAUROC = +0.0350 ± 0.1600, with A1 exceeding A0 in 2 of 5 seeds). We report this stochastic sensitivity as an empirical characterization of small-cohort EEG self-supervised learning rather than as a caveat to be minimized. Controlled comparisons against an explicit temporal-continuity objective, a SoftCLT-compatible soft instance-wise formulation, and three temporal kernels further show that the size and direction of the in-domain effect depend on how temporal information enters the objective. A representation-level diagnostic indicates that the canonical gain is not explained by increased effective rank. Finally, the compact encoder retains 0.9991 ± 0.0004 cosine similarity between FP32 and INT8 encoder outputs while achieving a 71% reduction in model size, demonstrating a practical consequence of the compact representation for resource-constrained physiological monitoring. Together, these results characterize patient-level temporal structure as a meaningful inductive bias whose in-domain downstream benefit is stochastic in small cohorts, while demonstrating measurable cross-cohort transfer of the learned representation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.