When Evidence Changes, Does Reasoning Follow? Revising Language Model Conclusions Under Changing Evidence
Abstract
Reasoning establishes relationships between premises and interpretations, which may be reconsidered when the premises change. We study how language models respond when relationships change while maintaining earlier reasoning in context. We change a fact that affects the correct answer while leaving the earlier reasoning in the context, then test whether models use the updated evidence or repeat an outdated conclusion. The experiments cover controlled arithmetic, logic, and navigation tasks, GSM8K, FinQA, CRUXEval, and Wikidata problems, and the STALEFLOW tool-action test, which examines whether updated information governs a subsequent action. We compare retaining the full history with dependency-based pruning: removing reasoning steps that depend on a changed premise, including their downstream conclusions, while preserving unaffected reasoning. This comparison tests whether stale reasoning obstructs revision and whether selectively removing it improves accuracy. Across these studies, we evaluate Qwen3-1.7B, 4B, 8B, 14B, and 32B; Qwen2.5-3B, 7B, and 72B; Llama-3.1-8B and 70B; Phi-4; Mistral-7B-Instruct-v0.3; and OLMo-2-7B, including base checkpoints, together with five OpenAI models. On the controlled problems, removing only the invalidated steps lowers Qwen3-4B-Thinking's rate of repeating the outdated answer from 0.969 to 0.070, and under immediate answering the effect appears in all 17 models tested and in all 40 model–dataset cells on the benchmark problems. A correction notice fails when models answer immediately but works when they reason first, so the better remedy depends on how the model answers. Across 1,800 multi-turn sessions, dependency-based pruning improves accuracy by 27.3 percentage points compared with retaining the full history. In STALEFLOW, Qwen2.5-7B uses an outdated value in 62.5% of immediate-action cases. These findings frame revision as maintaining relationships between evidence and conclusions: checking what has lost support, preserving what remains valid, and computing a new answer where needed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.