acceptodds
Under review as a conference paper at ICLR 2027

Sequence Composition in EHR Foundation Model Pretraining

Abstract

Autoregressive electronic health record (EHR) foundation models predict the next token in sequences of structured medical events. A common pretraining strategy concatenates patient histories into a global token stream and samples fixed-length windows from it. This gives patients with longer histories more training exposure and, under standard causal attention, allows tokens to attend across patient boundaries. These choices can misalign pretraining with downstream clinical prediction, where each prediction is based on a single patient’s history. We therefore investigate how the sampling and assembly of patient histories into pretraining sequences affect clinical prediction performance. We use the conventional global-stream approach as a baseline and compare alternative sequence composition methods that control how patient histories contribute to pretraining sequences. Across four MIMIC-IV v2.2 tasks, our best method improves macro AUROC, AUPRC, F1, and balanced accuracy, including relative gains of 9.3% in F1 and 2.7% in AUPRC under a fixed training recipe. We show that these gains extend across four additional dataset and tokenization settings, including INSPIRE and eICU-CRD, with higher macro AUROC, AUPRC, and F1 than the corresponding baselines in each case. Our findings show that restructuring existing EHR data can improve clinical prediction without changing the model or collecting additional data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.