What Does New-Task Accuracy Measure in Continual Learning?
Abstract
Does lower new-task accuracy indicate a reduced ability to learn? In continual learning, the answer depends on the update procedure and prediction problem. We examine this inference by pairing actual learning with current-only adaptation from the same inherited parameters, then changing prediction candidates at fixed weights. Across ten paired CIFAR-10 seeds, experience replay trails SGD by 11.151 percentage points in actual new-task accuracy yet leads by 0.971 points after matched adaptation. Restricting actual predictions to current classes reverses the mean ranking without changing the probe scores. The apparent deficit thus coexists with better within-task discrimination and better matched adaptation. Retention gives the comparison its practical meaning: both standard probes have zero old-task accuracy, while a 20% retention requirement favors actual replay over the tested validation-selected probes on CIFAR-10.1/.2. Fixed-history continuations and a full-budget CIFAR-100 ResNet experiment show where the rankings change; in the latter, ER leads on both actual and probe accuracy in all five seeds. Together, these results distinguish an observed new-task deficit from weaker adaptation of inherited parameters and show why current and retained performance must be evaluated jointly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.