Remembering What Changes the Future: Learning to Compress Agent Histories for Long-Horizon Tasks
Abstract
Long-horizon agents interleave reasoning, actions, and observations, causing their interaction histories to grow until context budgets force compression. Summarization and truncation can reduce this cost, but a coherent account of the past does not reveal which facts, goals, or constraints remain necessary for future control. The central question is therefore whether a summary preserves information that can affect subsequent decisions, rather than whether it provides a fluent account of past events. We introduce Policy-Aware Compression through Counterfactual Evaluation (PACE), a reinforcement-learning framework for learning future-relevant memory. At each compaction boundary, PACE compares continuations from the same environment state: one sees the full history and the other sees a candidate compact state. The one-sided difference in downstream utility gives a direct learning signal for harmful omissions, while task-state and safety signals preserve information that may matter beyond the immediate answer. The compactor is trained using candidate groups under hard memory budgets. A hybrid-execution analysis relates local future sufficiency to repeated compaction. In 8K single-boundary evaluations on HotpotQA and MuSiQue, PACE preserves substantially more downstream QA performance than prompt-based and learned compression baselines at the same compact-memory budget. The results support a simple principle: effective agent memory should prioritize what can change the future, rather than information that merely supports a coherent account of the past.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.