acceptodds
Under review as a conference paper at ICLR 2027

Answer Error Forgetting: Disentangling Memory from Temporal Reasoning over Continual Knowledge Streams

Abstract

Long-term memory is essential for language models operating over continual knowledge streams, where information is presented sequentially as knowledge accumulates, is updated, or is superseded over time. Recent long-term memory benchmarks extend beyond factual recall to evaluate capabilities such as knowledge updating and temporal reasoning over long histories. However, these capabilities are typically assessed through end-to-end question answering, making it difficult to determine whether an incorrect answer reflects a memory-related failure or insufficient temporal reasoning. We introduce DiMeR, a diagnostic benchmark that presents knowledge units sequentially and evaluates the same temporal questions under controlled evidence conditions. DiMeR contains 1,956 temporal question–answer pairs constructed from 674 Wikipedia articles, each associated with source-linked gold evidence in the corresponding knowledge stream. We further divide the questions into Easy and Hard groups according to whether the temporal cues expressed in the question are explicitly present in the knowledge stream. For each question, we construct five core evaluation conditions that vary when and how the supporting evidence is available. To more directly separate memory-dependent degradation from reasoning limitations, we additionally evaluate memory conditions on questions that the same model answers correctly when the relevant evidence is provided directly. Beyond answer accuracy, we measure evidence completeness and introduce Refresh-4K, which reintroduces the original gold evidence after continued knowledge ingestion. Experiments show that Hard questions are more sensitive to continued ingestion and that similar answer accuracy can correspond to different levels of evidence completeness. Evidence reintroduction also does not always recover the correct answer, indicating that answer correctness alone is insufficient to characterize long-term memory performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.