NoisyUpdateBench: Diagnosing Risk-Sensitive Revision in Long-Term Memory Systems
Abstract
Long-term memory benchmarks increasingly test knowledge updates, stale beliefs, and historical recall. Yet they rarely isolate two distinct questions: what state a system believes, and which revision action is warranted when applications assign different costs to false supersession, missed updates, and deferral. We introduce NoisyUpdateBench, a controlled diagnostic benchmark of noisy and retracting state streams with latent ground truth, versioned current and historical queries, and evidence-identical risk clones that hold observations fixed while varying only application costs. We formalize memory revision as utility-conditioned sequential decision making and evaluate transparent controllers that separate belief estimation, expected-risk action selection, and candidate quarantine. On a precommitted synthetic held-out suite of 120 base streams, dynamic Bayesian filtering with expected-risk decisions achieved the lowest pooled normalized risk cost (NRC) among ten configurations (0.676), improving over last-write-wins (ΔNRC = −0.655, 95% paired CI [−0.833, −0.482]); for risk-aware controllers, the clones preserved identical beliefs while inducing cost-dependent actions. On a sealed, independently authored synthetic external suite of 60 cases, the same controller again improved over last-write-wins and its matched fixed-policy variant, while comparisons with strong weighted-vote baselines remained inconclusive. In raw-language Track B, the same decision layer with automatic extraction reduced NRC relative to naive overwrite, full-context prompting, and Mem0 (ΔNRC = −0.770, −0.698, and −1.241; all paired CIs excluded zero) and achieved 0.962 exact-match extraction F1, while its NRC gap to a gold-extraction reference remained inconclusive. Scenario winners and quarantine effects varied across datasets, precluding universal-superiority claims. These results position NoisyUpdateBench as a diagnostic framework for utility-conditioned memory revision rather than another test of update accuracy alone
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.