Learning to Manage Context: Training Agents with Meta-Tools Beyond a Single Context Window
Abstract
An agent that continually accumulates interaction history eventually exhausts its context window, thereby limiting the task horizons over which they can be trained. We introduce the Concurrent Agent Rollout Process (CARP), which brings context-management mechanisms into a common framework for agent execution and training. CARP encapsulates these mechanisms as meta-tools, allowing agents to replace accumulated histories with summaries or delegate work to other agent instances. By retaining each generated token’s conditioning context, we can compute its log-probability for gradient-based training without requiring the complete execution to fit within one context window. We address obstacles which prevent agents from reasoning about context capacity. Using Group Relative Policy Optimisation (GRPO), agents learn to solve synthetic counting tasks whose data exceeds their context window, making context management necessary. Different reward and prompt conditions produce different strategies, including parallel task decomposition when the reward accounts for parallel computation savings. These results lay the groundwork for end-to-end training of agents on tasks whose execution extends beyond a fixed context budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.