Dynamic Sparse Training for sustained Plasticity under Non-Stationary Data
Abstract
Deep neural networks trained on non-stationary data, while retaining previously observed samples, represent a continual learning (CL) setting in which loss of plasticity remains an important failure mode, and catastrophic forgetting is not the primary challenge. Plasticity is the ability of a network to continue learning effectively from new data; its loss shows up as warm-started networks generalizing worse than networks trained from scratch on the same data. Weight sparsity has been used in CL to reduce task interference and mitigate catastrophic forgetting, but its role in preserving plasticity remains less understood, particularly when previously observed data remains available. We study whether unstructured sparse training (static sparse training (SST), gradual magnitude pruning(GMP), and dynamic sparse training (DST)) can mitigate plasticity loss in dataincremental learning settings. Across sample-incremental and class-incremental continual visual learning on CIFAR-10, CIFAR-100, and Tiny ImageNet, DST, in particular Rigging the Lottery (RigL), at moderate sparsities outperforms dense baselines, while in continual language fine tuning on TRACE, RigL at moderate sparsity likewise improves over dense training, with lower sparsity providing no advantage. RigL also compares favorably with dense training at most fractions of retained past data. Treating DST as an alternative training procedure rather than another plasticity intervention, we combine it with other plasticity improving interventions; both perform better on top of RigL than on top of dense training in continual visual learning and continual language fine tuning. Overall, our results suggest that dynamically evolving a sparse topology can help sustain plasticity across architectures and learning domains while remaining compatible with existing interventions for plasticity loss
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.