acceptodds
Under review as a conference paper at ICLR 2027

When Memory Is Enough to Decide but Not Enough to Reconsider: Memory Selection for Argument Survival

Abstract

Language models answer from selected or compressed memories, commonly evaluated by whether they preserve the current answer. Evidence redundant for that answer can become essential after a source is withdrawn. We study this gap across five evidence and reasoning benchmarks using budgeted memories and several open-weight language models. Memories that preserve a correct, supported initial answer can lose every evaluated revision basis, a set of records sufficient to justify the answer after withdrawal, even while the full history remains sufficient; this loss reaches 42.2% of such answers. Two histories can share a memory yet warrant different answers after the same withdrawal, so memory alone cannot identify the warranted answer. Paired interventions link model responses to argument completeness and dependence on withdrawn evidence. Replacing one necessary premise reduces revised-answer accuracy by 66.0 percentage points at matched record counts; abstention increases when alternative arguments share the withdrawn premise. These findings motivate argument survival as a selection objective across a prespecified family of withdrawals. Two records can each have zero marginal gain under this objective yet positive joint gain by completing a surviving argument. Pair-rescue implements simple joint admission to capture this value. Executable reasoning tasks and held-out GO-Plus ontology queries with surviving alternatives in the candidate pool show gains in retention and grounded review (correct revised answers with sufficient valid citations) at equal budgets, with current accuracy within a two-point non-inferiority margin. In a separate experiment, identifying and showing the model a verified basis already in memory adds 28.5 points of grounded review. A memory can be enough to decide yet not enough to reconsider.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.