WhenLoss: Diagnosing LLM Memory Beyond Answer Accuracy
Abstract
Conversation histories average 121K tokens in our LongMemEval setting, yet the default persistent store holds 5K tokens and must be written before future questions arrive. A wrong answer can reflect omitted evidence, missed retrieval, or limits of the reader, the answering model. We present WhenLoss, a controlled diagnostic protocol that compares truncated history, oracle evidence, complete stored memory, and retrieved memory while fixing the reader. These comparisons expose performance changes at intermediate interfaces; an analytic counterexample shows why their score gaps do not identify information loss. Across 500 LongMemEval questions, four of six baselines have write-side gaps exceeding retrieval-side gaps by more than 0.02 for each of three readers. This observation motivates inspecting stored evidence and improving its selection. Expected Predictive Compression (EPC) uses prospective questions to select supporting evidence before the actual question arrives. At a 5K-token write budget, EPC raises complete-memory Contains Match from 0.44 to 0.49 and retrieved-memory Contains Match from 0.38 to 0.42 over LLM summarization, averaged across three readers. Its complete-memory advantage increases to 0.11 at 2K tokens. These gains support the writing intervention under the tested budgets; evidence-retention checks and cost comparisons further characterize its benefits and requirements.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.