acceptodds
Under review as a conference paper at ICLR 2027

NGM: A Plug-and-Play Training-Free Memory Module for LLMs

Abstract

Recent studies have introduced memory-augmented architectures that decouple knowledge storage from neural computation, enabling more direct access to stored representations. However, existing approaches often rely on learned memory embeddings or explicit memory and retrieval components, requiring additional training or increasing architectural complexity. To address these limitations, we propose N-gram Memory (NGM), a training-free, plug-and-play module that repurposes a backbone model's pretrained token embedding space as memory. NGM consists of a Causal N-Gram Encoder and a Cosine-Gated Memory Injector. The former constructs N-gram representations by directly aggregating pretrained token embeddings, while the latter uses cosine similarity with ReLU gating to adaptively control the amount of retrieved memory injected into contextual representations. Unlike conventional memory-augmented architectures, NGM requires neither additional embedding training nor a separately learned memory bank or external retrieval infrastructure. We evaluate NGM on ten benchmarks across six language-model backbones spanning two model families and multiple scales, including four Qwen3 models (0.6B, 1.7B, 4B, and 8B) and two Llama-3.2-Instruct models (1B and 3B). NGM consistently improves the ten-task average across all six backbones. Notably, it improves HumanEval by 4.26, HumanEval+ by 2.41, and GPQA-Diamond by 6.56 percentage points. Furthermore, NGM extends beyond text-only language models to vision-language models, with Qwen3-VL-2B achieving improvements across all five VLM benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.