acceptodds
Under review as a conference paper at ICLR 2027

Lesson or Noise? Do Language Models Know Which Corrections Belong in Memory?

Abstract

Persistent memory requires assistants to decide how far a user's correction should extend beyond the current task. We study this decision in SCOPE, a controlled protocol in which a fixed coding answer is followed by corrections that vary the error, stated scope and tone, before the model generates its own persistent notes. Across fifteen models and up to five memory requests, whether a correction becomes a standing rule depends strongly on how the write is elicited. Under the default request, Claude Sonnet 5 writes a standing rule in 39 of 40 sessions even when the user explicitly calls the typo a one-off. Rephrasing the request as a decision reduces these writes to 7 while retaining the stated codebase convention in all 40 sessions. In exploratory tests, eliciting a recurrence judgment before writing has opposite effects across models: it reduces Sonnet 5's rules after a plain typo correction from 40 to 12, but increases them in nine of the ten other models tested. GPT-6 Luna calls the typo a one-off in all 40 sessions and then writes a rule in 26. An exasperated correction (“This doesn't even run … Ugh.”) that asks for no change in future behavior also raises Haiku 4.5's typo rules from 3 to 38. These findings identify memory writing as a request-sensitive decision: verbal judgments alone do not reliably constrain the rules that persist. We release the protocol and all model transcripts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.