Learning To Compress Context Into Reusable Experience Memory
Abstract
Historical examples and execution experience provide useful context for language models, but repeatedly processing their full text increases input length, prefill computation, and key/value (KV) cache requirements. We propose a context compression framework that converts retrieved experience into reusable experience memory for a frozen language model. The framework first induces common rules from training examples, then jointly encodes each rule and its associated example into a small set of layer-wise KV states. Segmented attention routes information through these memory positions, while compressor-side LoRA, task supervision, and distillation from a full-context view train the compressor to retain information relevant to answer or action prediction. At inference time, the frozen model reads the selected memory alongside the current input, separating experience encoding from online use. Experiments with Qwen2.5-7B-Instruct and LLaMA-3-8B-Instruct on ALFWorld and ContractNLI show improvements over textual retrieval baselines and competitive performance relative to parameter-efficient fine-tuning. In the ContractNLI efficiency evaluation, with caches precomputed and resident on the GPU, 15 memory positions replace approximately 1,347 experience tokens on average, yielding a \(4.83\times\) online speedup over processing the original context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.