Dual-Pathway Latent Memory Injection for Frozen Language Models
Abstract
Frozen large language models (LLMs) are usually augmented with external evidence by prepending the supporting passage to the prompt. This in-context injection is simple and effective, but prefill cost grows with evidence length, and the evidence competes with the generated history for attention. Recent latent injection methods instead write evidence into the KV cache, which integrates it more directly and at lower cost, yet these methods have not matched the accuracy of in-context injection. We explain this gap through the mechanics of factual recall: transformer LLMs rely on both MLP-encoded factual associations and attention-based attribute extraction, whereas KV injection acts only on attention, so under compression it preserves broad semantics but loses the fine-grained details needed for verbalization. We propose LAMP (Latent Augmented Memory with Parallel Pathways), which compresses evidence into a fixed-size latent memory and injects it into a frozen LLM through two complementary pathways: a KV-cache pathway that broadcasts compressed semantics across layers, and a residual pathway in which the LLM's hidden states query token-level evidence features to recover missing details. On eight fictional, real-world, counterfactual, and long-tail question-answering benchmarks, LAMP outperforms in-context injection overall and also scores above an in-context baseline fine-tuned on the same data. Because each memory holds one passage, longer evidence is handled by composing memories, which keeps accuracy stable up to 256K tokens at constant memory per query. LAMP also supports mid-generation updates, adding new evidence without re-processing the preceding context and correcting earlier answers on the fly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.