acceptodds
Under review as a conference paper at ICLR 2027

How knowledge flows from memorization to utilization in LLMs

Abstract

Humans can apply knowledge they have used before to new problems. Reasoning by analogy, they can also learn from this experience how to use other knowledge of the same kind, even knowledge they have memorized but never used. Do language models show the same forms of generalization? To study how models learn new knowledge without contamination from their pretraining data, we construct a synthetic world of fictional people, institutions, cities, and regions, described by 409 knowledge items that no model has seen before. We inject most of this knowledge through continued pretraining (CPT), teach models to reason with it through post-training, and track each knowledge item from injection to reasoning. We say that a model utilizes knowledge when it explicitly recalls the knowledge from memory as a step of its reasoning. We distinguish question generalization, utilizing knowledge seen in training questions to answer new questions, from knowledge generalization, utilizing knowledge that the model has memorized but never seen used in reasoning. Across four model backbones, CPT reaches 99.7% knowledge probe accuracy, yet CPT-only models explicitly recall less than 2% of the injected knowledge that reasoning questions require. Memorizing knowledge is thus not enough: models must also learn to recall internal knowledge during reasoning. Once models learn to recall, both forms of generalization emerge but remain limited. Recall also extends to knowledge the model was never taught, producing hallucination. We further propose Contextualized Utilization Encoding (CUE), which presents knowledge during CPT with diverse recall-style prefixes. Across four backbones and two synthetic worlds, CUE consistently improves knowledge generalization and also improves question generalization in three of four backbones. Our findings show that memorizing knowledge and learning to reason with it are distinct stages, and that how knowledge is injected shapes how it generalizes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.