Sorted by Date: Knowledge-Update Scores Cannot Tell Dates from Position
Abstract
Knowledge-update benchmarks are used to compare assistants and measure the benefit of memory updates. We show that a routine choice in evaluation code can undermine both conclusions: presenting evidence chronologically places the current statement after the superseded one. Time reading and position following then earn the same score, even when averaged over time-ordered layouts. We identify this confound in three update settings across six audited memory benchmarks. Paired exchanges of time labels or positions preserve input length and format, yet expose opposite cue preferences among eleven reader settings with similar native scores. Four LongMemEval readers differ by just 1.4 percentage points natively but by 16.9 points when averaged over five random orders; the lowest native scorer ranks first. Placing the most relevant evidence last puts the superseded session last on 27-44% of questions; GPT-4o-mini loses 23 points overall with lexical retrieval. Even an oracle update note shows no significant native gain across five readers. Its gain is 11.7 points larger when the superseded session is last (95% CI [6.3, 17.1]; exploratory paired analysis). Four memory systems extend the diagnosis to arrival order and stored dates. We recommend retaining native accuracy, adding random-order accuracy split by the competing pair's orientation, and using paired exchanges as diagnostics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.