Auditing Forecast Outputs Under Cohort Shift: Residual Learning Retains Conversion-Ranking Information Where Predicted Change Does Not
Abstract
A probabilistic forecaster produces several scores for each subject: its predicted change, its predicted residual spread, and its variability across model fits. These scores need not remain equally informative when the model is applied to a new cohort. We study this question in Alzheimer's disease (AD), where cortical thinning tracks disease progression. Here, we train a probabilistic forecaster on longitudinal Magnetic Resonance Imaging (MRI) data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) study to predict future regional brain cortical thickness, without using future-conversion diagnoses. We then freeze it and evaluate, on the Australian Imaging, Biomarker and Lifestyle (AIBL) data, how well each of three scores identifies subjects whose clinical diagnosis will worsen within 24 months (converters): predicted-change magnitude \(s_\mu\), learned residual scale \(s_\sigma\), and model variability \(s_PW\). As an in-domain reference, we evaluate a separate volumetric forecasting analysis within ADNI data, without cohort shift: on 127 held-out subjects (26 converters), its predicted change is informative (ROC-AUC ), residual scale is weaker (), and model variability is near chance (). In contrast, when the cortical-thickness forecaster is transferred from ADNI to AIBL 139 subjects; 14 converters), residual scale identifies converters with ROC-AUC , predicted change performs near chance (), and model-variability scores do not exceed chance (–). The advantage of residual scale over predicted change on AIBL is consistent across training seeds and leave-one-converter-out checks; its 95% interval includes zero at 24 months but excludes zero when the outcome window is extended to 36 and 54 months. We explain these differences theoretically: predicted change transfers only if its relationship to disease severity is preserved across cohorts; residual scale transfers if its distribution given severity is stable within each outcome class; and model variability reflects parameter sensitivity rather than disease state. These results argue for evaluating each forecast output separately under cohort shift; they do not show that uncertainty is generally superior, nor do they provide a calibrated clinical risk model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.