acceptodds
Under review as a conference paper at ICLR 2027

Mixture of Sub-Memories: Selective Updates for Continual KV-Cache Compaction

Abstract

Large language models can acquire knowledge from long contexts, but this approach becomes increasingly impractical in continual settings, where context accumulates over time: prefill cost grows quadratically with sequence length, and the KV cache grows linearly. Cartridges address this by distilling a long context into a small, fixed-size trainable KV cache. In deployment, however, knowledge is not static: new corpora arrive continually, and a memory with a fixed budget must be able to absorb each one while retaining what it already holds. We introduce two complementary approaches that extend training-free KV compaction updates built on attention matching that limit what each update overwrites: (1) a modular update, which routes each incoming corpus to one of several sub-memories under a likelihood gate and read one sub-memory per question, and (2) a monolithic update, which keeps a single memory but rewrites only a small, attention-selected subset of its slots. On LongHealth and QuALITY with two instruction-tuned models Llama-3.2-3B and Qwen3-4B, the modular update remains above the no-memory baseline at 16,384 slots and is at least as accurate as a text memory of equal size. Its gains on old phases shrinks with every update, and a controlled comparison attributes the loss to repeated recompaction without the source text rather than routing errors or memory size. The monolithic update improves over phases with Llama-3.2-3B and, at 8,192 slots, approaches in-context accuracy on LongHealth and reaches it on QuALITY, but collapses with Qwen3-4B. Our results show that a fixed-size KV memory can survive continual updates when writes are confined, and identify repeated recompaction as the main source of forgetting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.