Acquire, Then Retain: A Two-Stage Model of Memorization in Small Language Models
Abstract
The emergence of small agentic models makes the memorization capacity of small language models (SLMs) increasingly important to understand, yet it remains understudied relative to large-scale models. We inject synthetic facts into the training of five models spanning 3M–151M parameters and study memorization as a joint function of model scale, repetition count, exposure pattern, and time since exposure. We find that memorization at this scale is bimodal, i.e, a fact is either memorized or it is not, with little continuum in between. This motivates a two-stage retention kernel that separates whether a fact is acquired from how quickly it is forgotten. This two-stage kernel generalizes to a model size held out from fitting (, versus for a single-stage fit). On the axis of exposure, Spacing a fact's repetitions across training, rather than delivering them back-to-back, makes it harder to acquire at every model size, requiring up to twice as many repetitions. However, once a fact is firmly acquired, spacing slows its forgetting, extending its decay timescale by –. This benefit shrinks with scale though. Spacing thus trades acquisition for durability. Finally, facts acquired late in training, under small learning rates, are forgotten faster, with direct implications for how knowledge introduced during fine-tuning persists.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.