Learning What to Preserve: Parameter Attribution for Continual Learning in Large Language Models
Abstract
Large language models (LLMs) often suffer from catastrophic forgetting in continual learning: after learning a subsequent new task, they perform worse on earlier tasks. Existing replay and regularization based methods constrain model updates to mitigate forgetting, but may not reliably identify parameters that should be preserved or adapted. Common importance measures, such as Fisher information and gradient-based signals, capture local sensitivity rather than parameter contributions to model predictions. We propose a parameter-level attribution framework that traces output logits to individual parameters, enabling a mechanistic analysis of how the locations of important parameters overlap across tasks. We train LLMs independently on different tasks, analyze the resulting parameter importance patterns, and find that dissimilar tasks show lower overlap in LLMs layers. Motivated by this finding, we estimate parameter importance within each layer and selectively constrain updates to parameters important for previous tasks, while allowing less important parameters to adapt to new tasks. Experiments under both full fine-tuning and LoRA fine-tuning show that our method consistently outperforms all baselines across task settings on Llama and Qwen models, improving performance by at least 2 percentage points over replay-based methods and 3 percentage points over parameter importance based methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.