From Context Distillation to Latent Memory Management: Selecting and Activating Distilled Memories
Abstract
Context distillation internalizes a document into model parameters, allowing subsequent queries to be answered without repeatedly including the document in the prompt. While existing work primarily studies how to construct a memory from a single document, deploying a collection of such memories introduces two additional decisions: which memory is relevant to a query and whether the selected memory should be activated at all. We formulate these decisions as latent memory management with memory selection and memory activation, and introduce cache-compatible context distillation, which represents documents as independent LoRA adapters trained to operate on a query prefix KV cache produced by the frozen base model. This shared cache supports latent memory management without recomputing the query prefix, enabling an efficient system in which external retrieval produces a candidate set, internal routing selects one candidate, and Self-Gating decides whether to activate the selected memory. Across NarrativeQA, SQuAD, and LongBench, routing consistently improves over top-1 retrieval, while Self-Gating substantially recovers performance on context-agnostic queries with only small changes to document-specific QA. The same selection strategy also improves latent-token memories, suggesting that the management problem extends beyond LoRA-based representations. Our caching pipeline further reduces memory-construction time by up to 8.4×, peak memory by 4.1×, and training FLOPs by 21.9×, while adding little online management overhead. Overall, our results show that scaling context distillation requires not only writing memories, but also selecting when and which memories to use.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.