Address Migration for Replay-Free Continual Learning with Sparse Routed Memories
Abstract
Updating a representation can invalidate the coordinates used to retrieve previously learned memories. We study this problem in a continual learner that stores sparse routing prototypes and local experts, but retains no past examples or feature exemplars. Our approach, address migration, fits a linear map from new activations to the previous memory coordinates using paired evaluations of current unlabeled inputs. A wake–sleep schedule separates supervised expert updates from unsupervised projection updates, and a distance-profile regularizer limits geometric distortion during sleep. We characterize exact recovery for linear drift, separate finite-sample error from irreducible approximation error, bound changes in top- codes by the distribution of alignment error relative to coordinate margins, and show that cosine nearest-class-mean classification in code space is a special case of the architecture. On Split CIFAR-100 with a frontend pretrained on all classes without labels, migration with guided sleep achieves final accuracy, compared with for unguided migration and for the frozen configuration (three seeds), recovering about 80% of the accuracy cost of moving the projection and bringing a mobile representation within one point of the frozen configuration. Increasing routing capacity improves retention, whereas routing selectivity alone cannot compensate for weak per-task learning. The results establish coordinate alignment as a mechanism for keeping sparse routed memories addressable while the representation changes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.