Relational Window Stabilization for Repeatable Single-Window PPG Disease-Background Prediction
Abstract
Deep learning models for physiological signals are increasingly used for continuous monitoring of chronic disease background, yet their evaluation mainly relies on population-level discrimination metrics such as AUROC and AUPRC. In continuous monitoring, however, repeated measurements from the same individual should produce repeatable predictions when the underlying disease-background state remains unchanged over short intervals. Recent studies have revealed substantial prediction variability across repeated physiological measurements, highlighting a gap between discrimination performance and prediction repeatability. Here, we introduce Relational Window Stabilization, a training framework that learns repeatable single-window PPG disease-background predictions by exploiting relationships among multiple windows collected within the same monitoring episode. RWS uses multiple windows only during training, preserving individual-window supervision while reducing within-episode prediction variation without requiring inference-time aggregation. We evaluate RWS on the MC-MED PPG benchmark across nine chronic ICD-derived disease-background tasks using patient-disjoint chronological splits. Across different input durations, architectures, representations, and random seeds, RWS improves both discrimination and repeatability, increasing AUROC, AUPRC, and intraclass correlation coefficient while reducing prediction variability and diagnostic flip rates. These improvements remain under matched multi-window exposure controls and alternative repeatability-learning strategies. Cross-dataset evaluation on MIMIC-III further demonstrates that the repeatability advantage persists after checkpoint transfer and independent in-domain retraining. Together, these findings establish prediction repeatability as a complementary evaluation dimension for continuous physiological AI and demonstrate that same-episode relational supervision can improve single-window PPG predictions without changing the deployment paradigm
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.