Open-Book Pretraining for Language Models: Selective Memorization with Dynamic Retrieval
Abstract
Language models use their limited parametric memory to store facts that could instead be retrieved from an external store. We ask whether a model can decide, during pretraining, which facts to memorize and which to leave to a knowledge base (KB). We introduce Open-Book Pretraining (OBP), a pretraining method that makes such decisions dynamically by identifying positions where retrieved context reduces next-token loss, according to the current state of the model. Along with the pretrained language model, we learn a lightweight classifier that decides when to retrieve during inference. Using this approach, we train small language models (0.5B parameters) from scratch on Wikipedia. Evaluated with retrieval on held-out pages, a 360M OBP model reaches a lower perplexity than a 1.7B model trained on the same data, while retaining competitive perplexity even without retrieval. Relative to standard pretraining plus retrieval-augmented generation, OBP models use retrieved evidence better while memorizing fewer facts in their parameters. Our learned classifier is able to recognize when to retrieve during inference, adapting to different sizes of knowledge base. To our knowledge, OBP is the first framework that dynamically selects between parametric memory and external retrieval during pretraining, based on model- and KB-specific retrieval gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.