What Does a Model Learn from a Changing World?
Abstract
Models absorb world knowledge during pretraining, but they do not see a static snapshot of the world: pretraining data spans many time periods, showing facts change over time. In a controlled setting, we first ask whether models simply memorize each time period separately, or learn how knowledge evolves across them. We find the latter: compared to a model trained on a single time period, models trained on multiple periods (a) better preserve historically constant knowledge and (b) more efficiently learn updates that follow historical patterns. We then find that the order in which facts are presented determines how they are stored. When facts are presented chronologically, they are stored in individual slots, so an update changes only the facts it trains on. When facts from different time periods are instead shuffled together, models use class-based storage: facts that have changed together, such as athletes who switch teams together, share a representation, so an update to one of them spreads to the others. We find evidence of class-based storage in OLMo-2-7B, whose pretraining data, like our shuffled setting, mixes documents from different time periods. Beyond the known tendency of updates to spread among facts that share a relation and answer, we find correlational evidence that updates spread further in groups of facts whose answers changed more during pretraining. Our results show that models adapt to the changes they observe during training, and that the order in which they observe them shapes how, suggesting pretraining data ordering as a way to influence how models absorb new information.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.