: earning onstant-size atent ontext for Language Model Agents
Abstract
Language model agents have become highly capable across many areas. They accumulate increasingly long contexts of observations, actions, tool outputs, and intermediate reasoning. Without proper context management, agents' performance can degrade significantly as the context grows. Another major problem is that long context can drive costs up, making language model agents prohibitive in practice. In this work, we ask: can we effectively compress the agent context without affecting its performance? In this work, we propose : a framework for learning a constant-size latent context for a frozen language model agent. We leverage the strong representational power of the continuous latent token space: even as the rounds of agent interaction increase, the number of (latent) tokens representing the context remains constant. Concretely, we train a lightweight compressor to compress past context history into a constant number of latent tokens that we insert directly into the frozen language model's embedding sequence. We use a two-stage training procedure: behavioral supervised learning first teaches the compressor to preserve the full context, followed by reinforcement learning that directly optimizes the compressor using the downstream task reward. We argue that this training procedure can learn a compact latent representation of an agent's context, rather than relying on pure natural-language text summarization. Results on three widely used agentic benchmarks show that our method can improve performance up to 10% while reducing the token count to one-tenth of the original.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.