LANS: Tuning in Larger Null Spaces for Continual Learning
Abstract
Continual learning aims to enable models to sequentially acquire knowledge while maintaining a balance between stability and plasticity. To achieve this, null-space-based approaches enforce weight updates to lie in the null spaces of past input features, thereby preserving historical responses. Although this strict constraint offers a strong stability guarantee, it inevitably reduces plasticity, especially when model capacity is limited or the number of tasks grows. In this paper, we revisit the stability condition for the attention mechanism in modern Transformers. Unlike prior approaches that constrain each layer independently, we analyze how updates to the attention parameters affect both the attention affinities and the final output, and derive attention-specific stability conditions. These conditions define a feasible update subspace that is provably larger than the conventional null space, offering greater plasticity while preserving historical attention behavior to first order. Building on these results, we characterize the feasible subspace in closed form and develop LANS, a low-rank parameterization that efficiently learns weight increments within the derived constraints. Extensive experiments on multiple class-incremental learning benchmarks demonstrate consistent improvements over existing state-of-the-art approaches. Code is available at https://anonymous.4open.science/r/lans-207E.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.