MoNe: Modular Neural Memory for Efficient Long Context Inference
Abstract
Long-context reasoning is essential for applications such as personalized assistants, document QA, and agentic systems, yet standard in-context learning (ICL) becomes inefficient and unreliable as context length grows. We introduce (dular ural Memory), a lightweight plugin that equips frozen pretrained Transformers with test-time learned fast-weight memories, without modifying or retraining the backbone. MoNe processes long contexts in segments and writes information into these memories through block-localized gradient updates, without backpropagating through the Transformer backbone. During query inference, the adapted memories provide context-dependent residuals at each Transformer layer without including the original context in the model input, while remaining reusable across queries and incrementally extensible as new context arrives. Across RULER, BABILong, SQuAD, and HotpotQA, MoNe generalizes to 128K-token contexts while substantially outperforming the baselines at long context lengths and reducing both peak GPU memory and total FLOPs by approximately (80%) relative to ICL at 128K.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.