NewsCPT: News as a Substrate for Temporal Continued Pretraining
Abstract
Large language models (LLMs) become increasingly stale as the world changes beyond their pretraining cutoff. Continued pretraining (CPT) on newly produced text offers a natural way to update them, but it remains unclear which new facts such updates acquire, how far that knowledge propagates, and at what cost. We study these questions using news as a substrate for temporal CPT. We construct and release a globally deduplicated, timestamped corpus of 205M news articles from January 2024 to May 2026, along with dated question sets derived from it and Wikipedia's Current Events portal, and continue pretraining eight models from three families. Across three post-cutoff knowledge benchmarks, news CPT consistently outperforms both base models and matched-budget controls trained on post-cutoff FineWeb-Edu text. At the level of individual facts, accuracy improvement grows approximately log-linearly with how often the corpus states the fact, but those gains propagate only weakly to their consequences. The update degrades broader downstream performance, which replay with web data only partly mitigates. Mechanistically, retaining only the CPT-induced MLP updates preserves much of the new factual knowledge while improving several downstream losses, suggesting that acquisition and capability loss are partly separable. Our results provide a large-scale empirical account of how LLMs acquire new factual knowledge through CPT: factual acquisition is strongly associated with exposure in the update corpus, while propagation beyond the stated fact remains limited and acquisition and capability loss can be partially separated after training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.