Auditing Subject-Level Conformal Reliability in Medical Time Series
Abstract
Medical time-series (MedTS) classification often uses subject-independent evaluation, where models generalize to unseen patients who each contribute multiple potentially dependent observations. This creates a mismatch between the subject as the unit of deployment and the window as the unit commonly used for conformal calibration and evaluation. We conduct a systematic audit of subject-level conformal reliability in this setting. Across four EEG and ECG benchmarks, pooled evaluation can conceal substantial under coverage for individual unseen subjects; on ADFTD, near-nominal pooled and class-stratified coverage conceal the same failure. We show that pooled coverage forms a window-count-weighted average of subject-level coverage, while pooled calibration targets a quantile of a mixture of heterogeneous subject-specific score distributions. Across our diagnostic probes, training-time conformal optimization shows inconsistent effects on classification behavior, while similarity-weighted recalibration and selection-aware inference do not consistently resolve subject-level miscoverage. Established error-driven quantile adaptation can substantially improve reliability in sequential deployment when feedback is sufficient and representative, but its effectiveness depends on within-subject feedback volume and, for a shared-state controller, stream organization, and can reverse under prediction-dependent outcome observation. These findings concern continuous online adaptation rather than a threshold calibrated once and then fixed, and show that nominal pooled reliability need not reflect the reliability experienced by individual unseen subjects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.