acceptodds
Under review as a conference paper at ICLR 2027

Are LLMs Up-to-date on Medical Knowledge?

Abstract

Biomedical evidence continuously evolves as scientific and clinical findings can be revised, nullified, or reversed. However, large language models (LLMs) still largely rely on static parametric memory acquired during training. Evaluating whether LLMs retain obsolete medical knowledge then becomes critical, yet confounded by a attribution challenge: standard benchmarks cannot separate knowledge staleness from a failure to acquire current medical knowledge during pretraining. To tackle this challenge, we introduce MEDLAPSE, an evaluation dataset derived from 1,787 author-corrected health-science finding lineages of up to 20 version updates across 14,172 multi-version medRxiv preprints. Specifically, each superseded or outdated finding is matched to two elements: its updated successor and paper-specific retention claims, which then helps verify whether a staleness might stem from outdated memory rather than complete unfamiliarity with the paper. Across 16 leading open-weight and commercial LLMs, they consistently select the superseded finding in 30.3% of queries in Q&A testing on average while the retained, unchanged claims are correctly recalled. Notably, this staleness rate remains largely invariant across parameter scale, medical domain pretraining, and training cutoff dates. Furthermore, applying existing, post-hoc model correction methods reveals critical limitations: while “forget-then-update” approach designed based on unlearning algorithms such as NPO and RMU achieve high efficacy or correction rate (>94%) using the training prompts, their performance on paraphrased ones drops up to 6.1% below the baseline. These results demonstrate that current model correction methods induce prompt-specific token suppression rather than genuine factual revision, highlighting that safely deprecating obsolete parametric medical knowledge in LLMs remains a non-trivial but critical research challenge.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.