Sequence Composition in EHR Foundation Model Pretraining
Abstract
Autoregressive electronic health record (EHR) foundation models predict the next token in sequences of structured medical events. A common pretraining strategy concatenates patient histories into a global token stream and samples fixed-length windows from it. This gives patients with longer histories more training exposure and, under standard causal attention, allows tokens to attend across patient boundaries. These choices can misalign pretraining with downstream clinical prediction, where each prediction is based on a single patient’s history. We therefore investigate how the sampling and assembly of patient histories into pretraining sequences affect clinical prediction performance. We use the conventional global-stream approach as a baseline and compare alternative sequence composition methods that control how patient histories contribute to pretraining sequences. Across four MIMIC-IV v2.2 tasks, our best method improves macro AUROC, AUPRC, F1, and balanced accuracy, including relative gains of 9.3% in F1 and 2.7% in AUPRC under a fixed training recipe. We show that these gains extend across four additional dataset and tokenization settings, including INSPIRE and eICU-CRD, with higher macro AUROC, AUPRC, and F1 than the corresponding baselines in each case. Our findings show that restructuring existing EHR data can improve clinical prediction without changing the model or collecting additional data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.