Detecting Care: Treatment Policy and Retrospective Validation in Clinical Risk Models
Abstract
Clinical risk models are trained on outcomes recorded under a treatment policy, so a model can learn to detect care rather than risk. On a published ICU benchmark, masking one clinician-set feature reduces AUROC more than masking any of six non-manipulable physiological features. In this paper, we prove that a risk score with no intrinsic predictive value that triggers effective treatment has retrospective AUROC strictly below 0.5, and we derive a closed-form solution for the deficit as a function of treatment policy and treatment effect. We also show that this deficit does not vanish as intrinsic quality increases, and as a result, retrospective AUROC is not comparable across sites with differing treatment policies. For empirical evidence, we split 116 eICU hospitals by severity-adjusted vasopressor practice and find that a clinician-set feature which improves same-practice external validation degrades by a similar amount on transfer to a different practice. Moreover, the model's calibration error is concentrated on untreated patients, which is exactly the group these models exist to identify. The failure we characterize is invisible to discrimination, calibration, and same-practice external validation and our results can be extended to any acted-upon score, clinical or otherwise.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.