acceptodds
Under review as a conference paper at ICLR 2027

EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events

Abstract

Electronic health records (EHRs) encode clinical histories as (time, modality, code) tuples, whereas pretrained language models expect text tokens. Serializing each tuple as text inflates sequence length and repeatedly encodes the same structure. We introduce EHRAdapt, an adapter that maps tuples directly into a frozen language model's embedding space. Modality receives a learned embedding, time gaps enter through learned attention biases, and event codes receive dedicated vectors. Learning these event vectors is the central challenge: clinical vocabularies are long-tailed, leaving rare events with too few observations to estimate reliable representations independently. EHRAdapt therefore represents each event vector as the sum of a semantic prior and an evidence residual. The prior is a frozen embedding of the event's clinical description, from a biomedical language model trained on clinical ontologies, so it carries clinical knowledge. A shared learned projection maps this embedding into the model's input space, allowing the prior to supply clinical meaning even when observations are scarce. The residual is a learned low-rank, event-specific correction that refines this representation as evidence accumulates. We perform continued pretraining on records from about 4 million patients with three pretrained LLMs as backbones: OLMo2 1B, Llama3.2 1B, and OLMo2 7B. As the LLMs remain frozen, only the adapter is trained, which amounts to 0.1–0.6% of all parameters. The full adapter outperforms all ablations in held-out next-event prediction on every backbone. Removing the semantic pathway hurts rare events most, over ten times more than the most frequent ones, whereas removing the residual hurts overall prediction but improves it for the rarest events. On reportable infectious-disease and syndromic downstream classification tasks, EHRAdapt outperforms text-based LLM and count-based baselines, and both pathways improve rare-disease discrimination. The semantic prior and the evidence residual therefore play complementary roles, and these roles become visible only when results are broken down by event frequency rather than averaged.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.