Probabilistic Interpretation of Catastrophic Forgetting via Solution Density
Abstract
In continual learning, standard empirical risk minimization (ERM) can severely degrade a neural network's performance on earlier tasks, which is a phenomenon known as catastrophic forgetting. We study a solution-density interpretation of this behavior: what follows when current-task learning no longer favors low-risk solutions for a distal component of prior data beyond a reference density, and how much prior-task mass can be described in this way? Our framework combines a measure-theoretic risk decomposition with a proximal/distal split relative to the current task. We formulate a solution-density selection hypothesis that bounds the ERM learner's probability of selecting low-risk distal solutions by their normalized parameter volume. Under this hypothesis and label-permutation symmetry, we derive a chance-level lower bound on expected distal risk and a corresponding lower bound on final mixture risk. Within this framework, we further derive an inverse-factorial upper bound of on optimal solution density and selection probability, where is the number of classes with positive distal mass. CIFAR-100 proof-of-concept experiments illustrate the framework across class-disjoint, blurry, and generalized class-incremental settings through empirical risk-CDF comparisons and non-vacuous plug-in risk bounds.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.