A Retention Curve Certifies the Rehearsed Sentence, Not the Fact
Abstract
Language models are kept current by continual maintenance rather than retraining: new facts are written into the weights, the training sentences are replayed every round, and a stabiliser such as weight interpolation toward the pre-trained model limits drift. Retention is then measured on the replayed sentences themselves, on the premise that the item rehearsed every round is the item held best. We show that this measurement certifies a verbatim completion and nothing beyond it. On Qwen3-8B with 150 facts rehearsed for thirty rounds under a per-round weight interpolation, every training sentence is completed word for word every round, while the same fact in a held-out paraphrase is answered on .120 of the facts against .473 with weight interpolation off. The mechanism is the interaction of the two ingredients: weight interpolation alone lowers both measurements, rehearsal alone lowers neither, and together they rebuild exactly the training sentence. Even the training sentence, asked through the chat template, is answered on .087 of the facts against .560. The gap recurs on Mistral-7B, OLMo-2-7B and Qwen3-14B, on Wikidata facts and on 600 CounterFact and zsRE edits, under a periodic merge, and under later updates with no stabiliser. The standard rule for choosing a stabiliser strength, applied to the retention curve, picks a strength at which most of the paraphrase is already lost. At one distance from the pre-trained weights the retention curve scores every stabiliser family alike while the paraphrase separates them: on a float32 master an L-SP penalty answers the paraphrase on .393 of the facts where weight interpolation answers it on .140. And a per-round weight interpolation applied in place to bfloat16 weights delivers .02 of its nominal step by the thirtieth round, a shortfall the curve cannot show. We propose a release check that measures every maintained fact in a held-out paraphrase against the pre-trained baseline, holds some facts out of rehearsal, asks the training sentence through the chat template, and reports the stabiliser strength delivered rather than the one set.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.