acceptodds
Under review as a conference paper at ICLR 2027

Hindsight Memory-PRM: Supervising Memory Management with Auditable Hindsight Credit

Abstract

An LLM agent operating over a long horizon must decide, turn by turn, whether to write, merge, or skip memories, yet the quality of each operation is hard to judge when it is taken: whether a write was worthwhile becomes clear only much later, when the entry is retrieved and used to answer a question. We observe that memory trajectories already contain machine-readable evidence of this delayed utility—which entries were retrieved, which were cited by an answer, and which a controlled deletion test shows to be necessary for answering correctly. Hindsight Memory-PRM uses this evidence in two places: offline, to train an operation-conditioned memory-utility critic; and online, to settle intervention-calibrated, entry-level presence credit. The credit flows back along version chains to the operations that produced each entry, serving as an action-level proxy reward—without per-operation human labels and without replaying a full continuation for every action. On held-out LoCoMo, a local 8B policy reaches 77.5% under a fixed shared reader, above its API teacher (65.1%) and the strongest reproduced external configuration (74.7%) while using one eighth as many context tokens; it reaches 79.0% on LongMemEval. Controlled comparisons decompose the gain: observational feedback (entries being retrieved and cited) accounts for part of the improvement, and intervention-calibrated credit adds a substantial margin on top. Finally, the trained policy learns to consolidate related facts into multi-version entries, whereas open-loop controls given the same multi-version interface but no training signal do not recover the full improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.