When Is Forgetting a Capacity Problem? Dissecting Catastrophic Forgetting under Controlled Capacity Pressure
Abstract
Catastrophic forgetting is often attributed either to exhausted capacity or to optimisation interference despite sufficient capacity. The two explanations imply different remedies, and telling them apart calls for a measured capacity axis, which natural data do not provide: neither the stream's information content nor the model's effective capacity is directly known. To place forgetting on this axis, we build a continual-learning testbed from random facts, rules and noise, with known information content and calibrated model capacity. We define load as learnable information relative to measured capacity and vary it while holding each fact's training exposure fixed, which separates capacity limits from under-training. Three findings emerge in the testbed. First, storage of the old facts fails once they alone exceed capacity. Up to a load of 1 they are still stored almost in full, but as load grows a new linear readout needs more examples to decode them, and the model's own readout decodes them with narrower margins. Second, far below capacity, new facts in the old facts' format collapse old-fact recall within tens of steps, while less than 3% of their information is stored; training without new facts erodes recall much more slowly. During warm-up, collapse comes at a fixed cumulative learning rate, once the parameters have moved by about 1% of their norm; random perturbations of that size leave recall intact. Third, information survives the loss of recall and fades on separate clocks: the model's ranking of the correct answer outlasts exact recall, and brief retraining on a few old facts still restores others (a recoverable trace) once that ranking is at chance. During the collapse, unseen facts also become transiently hard to learn, while old ones stay easy to relearn. Among the consolidation methods we compare at three loads, utility-weighted replay retains the most useful facts and gives the highest future-task utility at each, and, at matched weighting, replay outperforms sleep-phase distillation with less compute. Weighting redirects retention toward useful facts even when all facts fit, but no method stores more than the measured plateau. A pretrained language model (Pythia-160M) fine-tuned on natural-language facts reproduces the collapse, its onset at a fixed cumulative learning rate and the recoverable trace; there, rendering the new facts as statements rather than in the old facts' question-answer format delays the collapse.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.