acceptodds
Under review as a conference paper at ICLR 2027

TaskFid: Benchmarking Agent Memory for Evolving Task-State Fidelity Across Threads and Sessions

Abstract

Long-term memory is essential for agents that collaborate with users across extended, multi-session tasks. Existing benchmarks evaluate retrieval, updating, and temporal reasoning, but rarely directly test evolving task-state fidelity: whether memory supports a correct, consistent, and actionable representation of an ongoing task. In long-horizon collaboration, task state is not a simple accumulation of history: previously valid information may be revised, superseded, or remain applicable only under specific conditions. Relevant history can therefore be retrieved without preserving task-state fidelity, leaving the supported task model outdated, inconsistent, or unable to support subsequent reasoning. We introduce TASKFID to evaluate evolving task-state fidelity across long-horizon task threads and sessions under isolated evolution, independent interleaving, and explicit cross-thread dependencies. At a frozen turn-level checkpoint, which we call an Anchor, independent queries evaluate two capability families: Task Reconstruction, including task-state and collaborative artifact reconstruction, and Resolution & Propagation, including dependency-conditioned actionability and scoped change-impact analysis. Longitudinal families revisit the same target across Anchors. Together, these evaluations measure whether memory supports recovery of and reasoning over the current task model throughout its evolution. Experiments across full-history conditions and representative memory systems show that preserving task-state fidelity remains challenging across task structures. Better retrieval does not consistently translate into higher fidelity, while longitudinal analysis reveals distinct update and retention errors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.