acceptodds
Under review as a conference paper at ICLR 2027

ScrubAgent: Context Scrubbing for Scalable Long-Horizon LLM Agents

Abstract

Large language model agents trained with reinforcement learning have achieved strong results on complex tasks requiring extended interaction. Recent work addresses the challenge of growing context by compressing or folding interaction history. These methods reduce context length but do not help agents locate critical information. Valuable discoveries remain mixed with failed attempts and repeated searches. Agents face disorganized history when making decisions. This motivates us to examine how humans process long content such as movies. Rather than passively receiving every frame, viewers actively construct structured memory. They note key plot points and skim over routine scenes. Important moments are highlighted and become easier to recall later. \ourmethod brings this approach to LLM agents. Through five operations, agents learn to compress routine steps, preserve important exploration and store discoveries as retrievable keyframes. We propose ScrubGRPO with process rewards that guide compression and memory decisions. On BrowseComp-Plus and SWE-Bench Verified, \ourmethod achieves 63.3% and 59.2% pass@1 respectively with over 65% context reduction. It outperforms full-context baselines and generalizes effectively to trajectories longer than those seen during training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.