acceptodds
Under review as a conference paper at ICLR 2027

Stress-testing prediction models: A systematic robustness evaluation on large-scale ICU data

Abstract

Limited robustness to real-world perturbations is a key source of generalization failure in healthcare, yet it remains underexplored in systematic evaluations. We introduce a robustness evaluation that assesses multiple model classes under systematized clinically motivated perturbations ranging from missingness, measurement noise, label errors, to distribution shifts. Using three large-scale intensive care databases (MIMIC-III, MIMIC-IV, and eICU), we evaluate linear models, tree-based ensembles, neural networks, and tabular foundation models on both in-hospital mortality classification and ICU length-of-stay regression. Across most perturbation settings, transformer-based models exhibit smaller performance degradation than tree-based methods. However, robustness varies substantially across perturbation types and model implementations, and tuned gradient-boosting methods remain competitive in several settings. We further characterize the computational cost of robustness, finding that the robustness gains of tabular foundation models are accompanied by substantially higher prediction time than optimized gradient-boosting models. These results highlight that robustness is multidimensional and model-dependent, and that model selection for clinical deployment should consider not only predictive performance under nominal conditions, but also sensitivity to realistic data perturbations and computational constraints. Our benchmark provides a systematic framework for evaluating these trade-offs in clinical tabular prediction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.