MemoryLifeBench: Evaluating Survival-Based and Utility-Ranked Memory Retention in LLM Assistants
Abstract
Long-running LLM assistants must decide which memories to retain under a storage budget. Predicting when a memory-related event occurs is a plausible intermediate objective, but does accurate prediction lead to effective retention? MemoryLifeBench examines this question using 10,152 right-censored records: synthetic records measure time to factual invalidation, whereas LoCoMo and LongMemEval records measure time to observed future evidence use. A Cox-trained predictor with 213,889 trainable parameters achieves a test C-index of 0.7218, exceeding five prompted and heuristic baselines. Nevertheless, MemoryLifeBench-TTL, which converts survival predictions into expiry thresholds, underperforms FIFO and reference-recency LRU on downstream question answering. MemoryLifeBench-Utility instead retains the top- memories by a separately supervised future-use score, increasing LoCoMo evidence retention from 0.6687 to 0.8765 at matched capacity. Its downstream improvement over FIFO and the TTL policy is exploratory: 80% of the 120-question LoCoMo sample comes from training conversations, and the 24-question validation/test subset does not establish superiority. An evidence oracle also outperforms retaining all memories with ordinary retrieval. These results separate event-time discrimination, budgeted retention, and retrieval quality, and motivate evaluating the memory decision directly rather than relying on a temporal prediction metric.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.