acceptodds
Under review as a conference paper at ICLR 2027

Turning Documents into Weight Updates with Paired Memory Tokens

Abstract

Large language models use new documents through prompting or fine-tuning, requiring either the document in context or additional parameter optimization. We introduce Paired Memory Adaptation (PMA), which learns to convert documents into reusable low-rank weight updates. PMA constructs separate read and write memories inside the frozen language model. Each read memory guides the construction of its paired write memory, and learned projection maps turn their states into update factors. We train the mechanism through context distillation: the adapted model matches a teacher that receives the document in context. New documents then produce updates without per-document gradient optimization. We evaluate fact recall and application in new situations on our Document Acquisition and Retention (DocAR) dataset, with answers assessed by a separate language model. After the same short, 600-step training on Qwen3-4B, PMA outperforms retrained Doc-to-LoRA (D2L) and SHINE on these tests with 50.5% and 19.0% fewer trainable parameters. The share of answers rated fully correct exceeds the unadapted model's by 21.65 percentage points on SQuAD (token F1 and a second judge agree; exact match does not). On most tasks, PMA's performance drops when the update comes from an unrelated document. Removing the learned readout reduces PMA's trainable parameters by 54.4% while retaining useful acquisition, although some public QA scores fall below the unadapted model's. Gains vary across tasks, and updates can disrupt answers to unrelated questions. In separate PMA and D2L experiments, training to preserve those answers reduces this disruption but also reduces gains on document questions. D2L's released checkpoint, trained far longer, beats PMA on nearly all tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.