Data Selection for Continual Learning: A Theoretical Analysis of Sample Scoring under Task Similarity
Abstract
In continual learning, a training example that helps a model learn a new task may also interfere with knowledge acquired earlier. Yet most theories of data selection assume a stationary learning objective, leaving unclear how sample prioritization affects learning and forgetting across sequential tasks. We develop a theory of score-based data selection in a structured linear teacher–student model, showing how feature and readout similarity between tasks determine a scoring rule's performance relative to uniform sampling. We analyze two representative scoring rules based on feature novelty and residual error. Novelty consistently improves the trade-off between learning the new task and retaining the old one, whereas residual-based selection depends on the relationship between the tasks: high-error samples help when the tasks are sufficiently aligned but increase forgetting when shared features support conflicting outputs. We further show that this qualitative ordering extends from score-proportional reweighting to hard budgeted selection, while tighter selection amplifies the effect. Experiments on CIFAR-100, CUB-200, and AwA2 qualitatively support the predicted dependence on task similarity and show that hard selection can produce larger retention effects than reweighting, while nonlinear CIFAR-100 experiments show a similar score ordering.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.