Cross-Modal Pretraining with Modality-Specific Risk Prediction from Polysomnography
Abstract
Polysomnography (PSG) records multiple physiological systems overnight and may contain prognostic information beyond sleep diagnosis. Using these signals for future disease prediction is challenging because PSG modalities differ in signal structure, availability, and relevance to individual outcomes. Although cross-modal learning can exploit shared physiology, merging all modalities into a single downstream representation may obscure information specific to individual physiological systems. We introduce a multimodal self-supervised framework that separates cross-modal pretraining from downstream risk integration. Independently parameterized encoders are trained using within-modality masked latent prediction and cross-modal prediction from other simultaneously observed signals. For time-to-event prediction, the encoders are frozen, and endpoint-specific whole-night aggregators and risk heads estimate risk separately from each available modality. These estimates are combined only at the prediction stage, allowing inference from incomplete modality sets without imputing missing signals or representations. We evaluate ten incident disease outcomes in 26,782 participants from three development cohorts and six harmonized outcomes in an external cohort of 1,000 participants. In pooled evaluation, combining PSG with basic clinical covariates achieves a mean Harrell’s C-index of 0.717. Although selected baseline performs better under the matched frozen probe, endpoint-specific downstream modeling increases the mean C-index of our PSG representation from 0.681 to 0.696 and yields higher point estimates than the SleepFM probe for eight of ten outcomes. Representation similarity exceeds the permutation null for 14 of 15 modality pairs, yet the predictive contribution of each modality varies across endpoints. Leave-one-cohort-out, external-cohort, and modality-removal evaluations further assess transfer under heterogeneous input configurations. These findings show that multi-modal PSG can benefit from cross-modal learning without requiring physiological modalities to share a single downstream representation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.