acceptodds
Under review as a conference paper at ICLR 2027

Scaling Gradient Projection for Continual Learning Over Long Task Sequences via Spectral Consolidation

Abstract

Gradient-projection methods address catastrophic forgetting in continual learning (CL) by retaining important activation subspaces and restricting subsequent parameter updates to minimally interfere with learned tasks. While these methods are competitive on standard CL benchmarks, their protection mechanisms become overly restrictive over long task sequences. We observe that in fixed-capacity models, protected subspaces formed by accumulating directions from successive tasks eventually saturate and stop admitting new directions. Even when constraints on parameter updates are relaxed, as in soft-projection variants, accumulated importance can ossify initially soft protection into effectively hard constraints. To address these limitations, we introduce Spectral Consolidation Gradient Projection Memory (SC-GPM), which combines activation usage across tasks in a bounded per-layer operator and recomputes the protected directions and their strengths after each task. We formulate the trade-off between protecting past tasks and reserving capacity for new learning as a convex optimization problem and derive an exact spectral solution. The solution reserves a specified amount of capacity for new learning by relaxing protection on less-used directions. Across DomainNet, OmniBenchmark, and ImageNet-1K streams of 100–160 tasks, SC-GPM improves final average accuracy by 2.18–3.64 percentage points over the strongest evaluated gradient-projection baseline on each dataset, with higher learning accuracy on all three. On domain-incremental Permuted MNIST with 200 tasks, SC-GPM outperforms the strongest evaluated gradient-projection baselines by 2.48 percentage points in final average accuracy and 1.24 points in accuracy averaged throughout learning (average anytime accuracy).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.