acceptodds
Under review as a conference paper at ICLR 2027

NoisyUpdateBench: Diagnosing Risk-Sensitive Revision in Long-Term Memory Systems

Abstract

Long-term memory benchmarks increasingly test knowledge updates, stale beliefs, and historical recall. Yet they rarely isolate two distinct questions: what state a system believes, and which revision action is warranted when applications assign different costs to false supersession, missed updates, and deferral. We introduce NoisyUpdateBench, a controlled diagnostic benchmark of noisy and retracting state streams with latent ground truth, versioned current and historical queries, and evidence-identical risk clones that hold observations fixed while varying only application costs. We formalize memory revision as utility-conditioned sequential decision making and evaluate transparent controllers that separate belief estimation, expected-risk action selection, and candidate quarantine. On a precommitted synthetic held-out suite of 120 base streams, dynamic Bayesian filtering with expected-risk decisions achieved the lowest pooled normalized risk cost (NRC) among ten configurations (0.676), improving over last-write-wins (ΔNRC = −0.655, 95% paired CI [−0.833, −0.482]); for risk-aware controllers, the clones preserved identical beliefs while inducing cost-dependent actions. On a sealed, independently authored synthetic external suite of 60 cases, the same controller again improved over last-write-wins and its matched fixed-policy variant, while comparisons with strong weighted-vote baselines remained inconclusive. In raw-language Track B, the same decision layer with automatic extraction reduced NRC relative to naive overwrite, full-context prompting, and Mem0 (ΔNRC = −0.770, −0.698, and −1.241; all paired CIs excluded zero) and achieved 0.962 exact-match extraction F1, while its NRC gap to a gold-extraction reference remained inconclusive. Scenario winners and quarantine effects varied across datasets, precluding universal-superiority claims. These results position NoisyUpdateBench as a diagnostic framework for utility-conditioned memory revision rather than another test of update accuracy alone

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.