Removing Sequentially Trained Modules: Effects on Retained Sources and Subsequent Learning
Abstract
We show that teaching a language model new information through separate, trainable components does not ensure that removing a component leaves what other components learned unaffected, and that removal can make subsequent learning easier. This matters for models that are continually learning from new datasets or environments while needing to dynamically disable some components and leave others activated. We test this by giving each dataset its own separately trainable component, called a module, while freezing the original model weights and varying which modules are active while subsequent modules learn. On synthetic tasks across five seeds we find that removing the first module trained reduces accuracy on the remaining datasets by 93.7 percentage points; when randomly disabling each earlier-trained module with probability .75 during training we avoid this observed degradation entirely at the cost of 58.3% more fitting steps. A separate three-seed experiment showed that removing four sources and their recorded dependent modules reduced the total steps needed to learn a new source to 43.6% of the no-removal baseline, although this removed all eight earlier modules.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.