Sparse Multi-Landmark Supervision for Longitudinal EHR Survival Prediction
Abstract
Electronic health records (EHRs) accumulate irregularly over years, yet survival models that support longitudinal risk prediction are commonly trained at only one prediction time per patient and then queried repeatedly as new clinical observations arrive. We study sparse multi-landmark supervision, which trains from several points along each patient trajectory, as a model- and objective-agnostic strategy to reduce this mismatch. Across eight tasks, two health systems, and six model families, training with three landmarks improves integrated Brier score in all 48 task-model configurations and concordance index in 45 of 48 under multi-landmark evaluation. These gains persist when optimizer updates or input tokens are matched, when recent clinical history is progressively truncated, and across retraining. Multi-landmark supervision also weakens coupling between predicted risk and EHR recording patterns (e.g., code density and redundancy). On EHRSHOT, three-landmark training can further compensate for substantially fewer training patients. Beyond predictive performance, we introduce trajectory-aware evaluation for models used repeatedly over time. Multi-landmark supervision makes successive predictions more temporally consistent, while direct trajectory analysis reveals that consistency and informative risk evolution are distinct properties. A ranking-augmented proof of principle shows that explicitly relating predictions across landmarks can alter trajectory direction. In a nutshell, our results show that better use of existing longitudinal records can substantially improve survival prediction, while motivating models and evaluation that explicitly account for how risk evolves within patients over time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.