In-Model Context Compression with Agentic Soft Token for Long-Horizon LLM Agents
Abstract
Long-horizon LLM agents play an important role in solving complex and sustained tasks, with ever-growing interaction contexts. However, these accumulated contexts raise substantial computational costs and may obscure key decision-relevant evidence. Existing methods typically train external compressors or context-management agents to rewrite context as shorter hard-token sequences, but such external selection may not align with the LLM's evolving representation and is necessarily lossy. In this paper, we introduce Agentic Soft-Token Compression (ASTC), a training-free compression method that uses soft tokens within LLM agents for long-horizon tasks. Specifically, ASTC uses the model's native recurrent dynamics to determine already absorbed information, leveraging the commonly used hybrid Gated Delta Network (GDN)–attention architecture. It then carries this information forward as agentic soft tokens while retaining exact values as hard tokens. The resulting compact sequence reduces later-layer computation and yields reusable KV caches, extending the savings across the agent's trajectory. Extensive experiments on long-horizon tasks show that ASTC matches or exceeds the uncompressed agent's accuracy with up to 50.68% fewer end-to-end Attention Pairs, while trained compressors require – more Attention Pairs per successful task. These findings highlight agentic soft tokens as a model-native way to carry absorbed context forward as the task representation evolves. Our code is available at https://anonymous.4open.science/r/agentic-soft-token-8e3b6a/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.