Towards Transferable Parameter Updates for Continual Learning in Large Language Models
Abstract
Continual learning (CL) for large language models (LLMs) aims to enable sequential knowledge acquisition while mitigating catastrophic forgetting. However, most existing LLM CL methods primarily regulate parameter modifications, such as determining which parameters should be updated or preserved, while rarely exploring how effective task-specific update behaviors can be learned and generated. This limitation hinders their ability to facilitate knowledge transfer across tasks. To address this challenge, we rethink parameter updates in CL from the perspective of learning adaptive update strategies. Specifically, we introduce Meta-HA, a meta-learning framework that learns transferable update strategies through Hyper-Adapter modules. Hyper-Adapter module consists of a meta-learner and a controller, where the meta-learner captures transferable update priors across tasks through outer-loop optimization, and the controller instantiates these priors to generate task-specific updates and consolidate new knowledge into memory during inner-loop adaptation. Extensive experiments on multiple CL benchmarks demonstrate that our method consistently outperforms state-of-the-art baselines in both average performance and backward transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.