acceptodds
Under review as a conference paper at ICLR 2027

Correcting Errors Is Not Enough:Exact Audits of Answer Preservation on Shared Tasks

Abstract

Memory-augmented large language models (LLMs) need to use helpful memory while preserving correct answers when memory is unreliable. We call the latter ability correct-answer preservation. Comparing this ability across model config- urations is difficult because different configurations may answer different tasks correctly without memory, while missing responses and incomplete task alignment create further uncertainty. We propose a shared-task audit that compares corruption rates on tasks that every configuration in a declared group answers correctly at baseline. When task alignment or outcomes are incomplete, we derive sharp finite- set bounds over all task assignments and missing-result completions consistent with the retained summaries. We also prove an exact solver that needs to enumerate missing-correct counts only for the two target configurations while preserving the same outer bounds as full search. The framework further identifies when additional overlap information is sufficient to resolve an otherwise uncertain comparison and when task-level labels are required. On reported MEMIMPACTBENCH counts, pair-specific shared-task support determines 14 of 15 harmful-memory directions, while one all-six shared set determines eight misleading-memory and seven stale- memory directions. We also validate the audit on public row-level StrategyQA conflict records spanning four model configurations and 12 incorrect-prior con- ditions: all 144 count-based intervals contain the matched-task truth, and all 43 strict direction certificates match the directly counted sign. These results show that rates computed on different task subsets do not necessarily support the same cross-model ordering. The shared-task audit states which preservation differences are supported by the available evidence and which require additional information.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.