RECAP: Training Long-horizon LLM Agents with Context Compaction
Abstract
Context compaction enables large language model (LLM) agents to tackle long-horizon tasks within limited context budgets by replacing interaction histories with compressed summaries. However, learning effective compaction jointly with the agent policy remains challenging, as each summary’s effect on task success unfolds over many subsequent interactions. We introduce RECAP, a reinforcement learning (RL) framework based on the principle that a summary should preserve the behavior induced by the original interaction history. We formalize this principle with a behavior distortion objective that measures the divergence between the agent’s subsequent behavior under compressed and original contexts. Its gradient decomposes into two complementary learning signals: distortion-based reward for summary generation and self-distillation for subsequent actions. Together with the task-success objective, these signals jointly optimize context compaction and the agent policy. Experiments on BrowseComp-Plus and SWE-bench Verified demonstrate the benefit of \method compared to other RL baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.