acceptodds
Under review as a conference paper at ICLR 2027

When Stale Facts Win: Why Language Models Lose Context in Corrections, and How to Measure, Predict, and Fix It

Abstract

If a context mentions a fact and later explicitly corrects it, a language model should respond with the new value. Often the old value remains in effect. We investigate this failure as a distinct ability to track updates and distinguish it from the mere retrieval of information. A causal audit first examines, for each model, whether attention-based sentence rankings are in fact necessary and sufficient for the response. A pattern emerges: the models focus on the correction but still select the outdated value. Using StaleBench, we measure this behavior across different context lengths, update types, and model sizes. The results show that update tracking neither reliably scales with model size nor is predicted by standard longcontext tests. A position swap also provides evidence that smaller models often follow a positional heuristic; for large models, this finding regarding the mechanism remains tentative. Common prompt and decoding methods do not robustly resolve the problem. We therefore evaluate a training-free context pruning method that removes outdated value-binding sentences before generation. It significantly improves the responses, provided that the conflicting statements are detected. It is this detection that proves to be a central bottleneck on natural Wikipedia text. This paper does not offer the familiar recommendation to delete old values, but rather a causally verified measurement and diagnostic procedure, a benchmark for update tracking, and an explicit analysis of the assumptions and limitations of context-based repairs. In practice, this implies that agent stores and retrieval systems should resolve conflicting versions prior to generation and test update tracking separately from information retrieval.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.