Beyond In-Task Accuracy: Evaluating Task-Specialized Memory under Relevance Shifts
Abstract
High accuracy on anticipated queries does not establish whether an LLM memory remains useful when different aspects of past experience become relevant. We isolate this evaluation gap by changing a memory writer's announced task while holding the history and decision rule fixed. In a controlled feature–reward task with exact scoring and reward-reversed histories, the writer is either promised that one feature dimension will matter or told that any dimension may matter. We freeze its memory and evaluate both announced and shifted queries. In the small task, cued memories from GPT-5.5 and GPT-5.4 achieve 100% in-task accuracy but 50.0–51.7% after a relevance shift. Uncued memories achieve 100% on both without exhausting the same 100-word allowance. Thus, identical in-task scores conceal different cross-query capabilities even when a broader useful record fits. The shift deliberately violates the cued writer's promise: specialization is appropriate for the announced task, not evidence of irrational forgetting. A separate 100-dimension task exposes model-dependent encoding and reading limitations rather than a uniform shift effect. A post-hoc retrieval audit distinguishes another source of error: all 302 non-tie lexical actions follow the selected reward sum, including 44 wrong full-history answers. Together, these results support evaluating anticipated-task accuracy, relevance-shift performance, and evidence use separately, rather than treating success on the anticipated task as a sufficient measure of memory quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.