Prediction Is Not Protection: What Models Forget During Context Compaction
Abstract
Long-running LLM agents replace old conversations with compact memory. Protecting selected facts can improve overall recall while making information left unprotected less reliable. We study this tension through a paired evaluation that links each fact's pre-compaction forecast, untreated omission, and outcome after key-only protection. With supplied importance labels and controlled fact budgets, gains in overall important-fact recall coexist with losses on the same unprotected facts under self-selected, random, and cue-based policies. At the widest memory-loop budget, external protection retains 76–93% of raw-log recall at 10–29% of its storage, but a simple repetition rule remains competitive; self-ranking is model-dependent. Naturalistic tests qualify these findings. LoCoMo establishes neither general cue superiority nor a best selector. On 12 Multi-Session Chat dialogues at approximately matched measured storage, total point gains occur without detectable unprotected-fact loss. Additional key-free controls do not establish a benefit from stronger wording; self-selection retains a nominal gain over plain compaction. These limited-source results do not establish the absence of collateral cost or a universal selector advantage. Prediction accuracy, overall recall, and reliability on the same unprotected facts therefore require separate evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.