acceptodds
Under review as a conference paper at ICLR 2027

REAL-EHR: Benchmarking Open-Loop Clinical Event Forecasting

Abstract

Open-loop EHR forecasting requires models to predict beyond the observed patient record without receiving future observations, yet most models are evaluated with teacher-forced next-event prediction. Teacher-forced performance may not reflect open-loop forecasting, and open-loop evaluation itself depends on how generated futures are scored. We introduce REAL-EHR, a benchmark that compares matched patients and targets while varying information access, scoring criteria, and rollout budget. We adapt 17 model architectures to a common forecasting interface. At a 32-rollout budget on MIMIC-IV, teacher-forced and estimated open-loop NLL rankings show little agreement (), with 62 of 136 pairwise model orders reversed. At the same budget, all 17 learned models have higher estimated marginal NLL than a frequency baseline, while 16 improve candidate coverage and window occurrence. These NLL comparisons are budget-sensitive: increasing the rollout budget to 512 moves nine of ten audited models below the frequency baseline by point estimate, while weak agreement with teacher forcing persists within this subset (). Model implementations and evaluation code are available at https://anonymous.4open.science/r/REAL-EHR-E237.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.