acceptodds
Under review as a conference paper at ICLR 2027

The Capacity Limits of Sequential Knowledge Editing

Abstract

Sequential knowledge editing updates a language model one fact at a time without retraining it from scratch. Most recent methods make each edit by changing the same feed-forward weights. This works well for shorter edit sequences, but performance can fall sharply as more updates accumulate. We study a simple alternative. Instead of writing every edit into the model weights, we keep the model frozen and store the optimized hidden value for each edit in an external memory. We call the resulting method MoKE. Several context-dependent keys are stored for each edit so that paraphrases can retrieve the same value. On the standard 3,000-edit protocol with GPT2-XL and Llama3-8B, MoKE gives the best reported result on 15 of 18 metrics. In a controlled comparison with DeltaEdit using the same code, machine, precision, seed, and edit stream, MoKE improves CounterFact specificity by 35.0 and 49.8 points. At 10,000 edits, MoKE retains 97.8 efficacy while state-of-the-art falls to 13.1, and this endpoint is replicated with a second seed. Keeping the base model frozen also preserves downstream accuracy much better after editing. External memory has clear costs. Memory grows linearly with the number of stored keys, the retrieval threshold must be calibrated, and specificity still falls as the memory grows. Direct measurements show that this loss is not explained by an increasing rate of paraphrases being matched to another edit. These results show that the scale of sequential editing depends strongly on where the edits are stored. Code: https://anonymous.4open.science/r/moke-editing-84FD/README.md

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.