Remember What Matters: Relevance-Banded Compression for Long-Horizon LLM Agents
Abstract
While long-horizon large language model (LLM) agents show immense promise in complex environments, long-horizon execution inevitably leads to a continuously growing interaction history, often triggering context explosion and performance degradation. To mitigate this issue, existing context compression methods typically compress historical context using uniform, blanket strategies. However, such indiscriminate operations fail to recognize the dynamic value of intermediate steps, risking over-compression of fine-grained, task-critical execution details. To address this limitation, we propose Relevance-Banded (ReBand) Compression, a training-free framework that uses semantic relevance between historical chunks and the agent's current runtime state to guide differentiated compression. Specifically, ReBand assigns unsummarized chunks to low-, medium-, and high-relevance bands, applying aggressive compression, concise factual summarization, and structured key-value retention, respectively. A single compressor call applies this guidance to update the rolling summary, while the latest interaction remains verbatim. Across three long-horizon task benchmarks, ReBand consistently outperforms ACON-UTCO, improving primary task metrics by 3.16–4.8 points while reducing peak context by up to 42.4%. These results demonstrate that runtime relevance provides a simple, transferable signal for allocating compression fidelity across context chunks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.