acceptodds
Under review as a conference paper at ICLR 2027

Did Your LLM Really Forget? Measuring Factual Recall During LLM Post-Training

Abstract

Post-training shapes LLM behavior, e.g., by aligning models with user preferences and scaling test-time compute through long reasoning trajectories. While post-training rarely teaches models new facts about the world, it might induce forgetting of factual knowledge acquired during pretraining. Previous studies of this phenomenon either rely on aggregate measures that mask individual facts or compare only two checkpoints. We propose a different way to approach this problem, investigating how the probability of recalling individual facts is altered across the entire post-training trajectory. Using repeated rollouts across intermediate checkpoints, we show that individual facts' recall probabilities fluctuate substantially during post-training, repeatedly increasing and decreasing even when aggregate accuracy remains relatively stable. Although changes in the underlying probability of recalling each fact largely explain the magnitude of these fluctuations, whether a fact is classified as forgotten or reappearing is driven primarily by sampling variability. Thus, we show that prior methods can substantially overestimate persistent knowledge loss. Finally, we formalize a lower bound to certify the knowledge loss in bits. We show that our metric is more consistent than existing ones and can be easily extended to account for recall fluctuations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.