Benchmarking Knowledge Acquisition with Causally-Related Synthetic Future Events
Abstract
Large language models rely on static pre-training corpora, resulting in outdated knowledge and motivating methods for post-hoc knowledge integration. Existing approaches for evaluating knowledge updates either suffer from rapid contamination or rely on counterfactual edits that conflict with a model's existing knowledge. In this work, we introduce CASCADE, a model pipeline that generates causally interconnected fictional future events while preserving local and global consistency across a large event graph. We use CASCADE to construct a benchmark spanning a taxonomy of question reasoning types, avoiding contamination while enabling controlled, fine-grained evaluation of knowledge acquisition and reasoning. Using CASCADE, we benchmark retrieval on proprietary and open-weight models and continued pretraining (CPT) on open-weight models, finding that retrieval consistently outperforms parametric knowledge injection and that reasoning across multiple linked events is substantially harder than reasoning on a single-event. We show that reasoning-focused reinforcement learning improves accuracy under retrieval, and that retrieval methods degrade less as the number of events increases compared to CPT. These results position CASCADE as a scalable, contamination-resistant testbed for knowledge updating and multi-event reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.