IncreMem: Compressing Long-Term Agent Memory in Increments of Information Gain
Abstract
Long-horizon LLM agents accumulate interaction histories whose length makes full-context attention prohibitively expensive. Latent memory compression reduces this cost by condensing the history into gist tokens that the model attends to instead of the raw context. Existing methods, however, compress at a uniform rate, as if information were spread evenly across tokens, so a multi-token fact is readily split across gist tokens while fillers and restatements take an equal share of the budget. We argue that capacity should instead follow the information each token contributes. Casting compression as a rate–distortion problem, we show that if a memory unit's distortion is a convex function of the information it holds, regardless of its length, spans of equal information gain minimize the distortion of the LLM's output. We measure this gain through a retrieval head, where a token is informative to the extent that it shifts attention away from its predecessor's focus and onto itself. The resulting method, IncreMem, closes a span at every fixed increment of cumulative gain, so that each gist token carries an equal share of information rather than of tokens. Across three long-context benchmarks and three backbones, IncreMem loses virtually no accuracy at half the memory of the uncompressed backbone and outperforms uniform gist compression and KV eviction in nearly every setting up to a compression ratio of 16. Adding only a lightweight gist attention module and gain estimator to a frozen LLM, it runs faster with 37.5% less peak memory than the uncompressed backbone at 128K tokens, within 3% of the latency of uniform gist compression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.