Rethinking Cross-Task Generalization in Continual Learning
Abstract
Continual learning (CL) aims to sequentially learn tasks with reduced forgetting and improved cross-task generalization. Although the two goals are usually treated separately in previous works, our theoretical analysis shows that improving cross-task generalization helps reduce forgetting. Previous works address the cross-task generalization problem by introducing large-scale reference data or freezing the pre-trained weights. However, the former is computationally expensive, while the latter does not assess the model's generalization ability in full fine-tuning. To improve the cross-task generalization with full parameter tuning, we find that parameters with largest (weight) updates in fine-tuning can be vulnerable in cross-task generalization and suffer more from forgetting, while parameters that have consistent updates in pre-training and fine-tuning are more robust. Inspired by this, we propose a double filtering mask to filter weight updates in fine-tuning that are inconsistent with pre-trained weight magnitudes. Although simple to implement, our method significantly improves model's cross-task generalization and achieves comparable performance to CL models that require large-scale reference data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.