Input Representations Meet Pretraining Objectives: A Controlled Study of Generative EHR Foundation Models
Abstract
Accurately modeling irregular event timing and numerical measurements is central to zero-shot prediction with generative Electronic Health Record (EHR) foundation models. Recent work has studied the impact of input event representations and loss function separately, overlooking their correlation and joint impact. Additionally, there is no comprehensive benchmark that systematically evaluates zero-shot inference performance on a large and diverse number of tasks. To address these problems, we conduct a control study that aligns the tokenizer, backbone, and training budget to isolate the effect of input embeddings and pretraining loss across three model families: Discretized, Continuous-fused, and marked temporal point-process (MTPP). For zero-shot evaluation, we introduce ZEHRO, a comprehensive zero-shot benchmark that spans two datasets, three task types, and two time horizons. Our results show that continuous representations and loss functions benefit regression, while discretization leads in classification. The point-process approach achieves the best performance in predicting large time intervals and remains competitive in long-horizon time-to-event prediction. Our ablation studies further show that changing the pretraining objective generally has a larger effect on performance than changing the input embeddings, and that these changes can have aligned or opposing effects across different tasks. These findings reveal task-dependent trade-offs in input representations and objective choices, informing their joint design in future generative EHR foundation models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.