Attention Guided Policy Learning for Context Selection in Long Horizon Agents
Abstract
Large language models increasingly serve as the reasoning core of agents that operate over long, multi-turn interactions. Because every new observation is ap- pended to a running context, the context such systems present to the model grows far faster than it is used, and an increasing share of it is irrelevant to the query at hand. Recent works show that this bloat is not a purely computational nuisance: accuracy degrades sharply and non-linearly as soon as distractors enters the con- text, with most of the damage done long before the window is full. Compressing or filtering context with signals external to the target model, such as an auxiliary scorer’s perplexity or embedding similarity, discards a resource the model already computes internally: a small set of attention heads and activation directions has been shown to track, with high fidelity, which tokens are being retrieved, copied, or bound into an answer. We formulate the problem of learning a context-curation policy that reads this internal, attention-level relevance signal directly from the ex- ecuting model and combines it with an outcome-grounded reinforcement-learning objective to select the subset of an agent’s accumulated context that should be forwarded to the next inference call. We propose CURATE, which learns this pol- icy with reinforcement learning, assigning credit for selection decisions over long horizons from sparse task-level reward. CURATE learns to identify and discard distractors without any relevance labels, retaining less than 10% of the context while preserving downstream accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.