acceptodds
Under review as a conference paper at ICLR 2027

RECAP: Training Long-horizon LLM Agents with Context Compaction

Abstract

Context compaction enables large language model (LLM) agents to tackle long-horizon tasks within limited context budgets by replacing interaction histories with compressed summaries. However, learning effective compaction jointly with the agent policy remains challenging, as each summary’s effect on task success unfolds over many subsequent interactions. We introduce RECAP, a reinforcement learning (RL) framework based on the principle that a summary should preserve the behavior induced by the original interaction history. We formalize this principle with a behavior distortion objective that measures the divergence between the agent’s subsequent behavior under compressed and original contexts. Its gradient decomposes into two complementary learning signals: distortion-based reward for summary generation and self-distillation for subsequent actions. Together with the task-success objective, these signals jointly optimize context compaction and the agent policy. Experiments on BrowseComp-Plus and SWE-bench Verified demonstrate the benefit of \method compared to other RL baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.