DualKL: Distributional Trust Region for Continual LoRA Adaptation
Abstract
Continual adaptation requires large language models to learn new tasks without losing capabilities acquired from earlier ones. Existing LoRA-based methods mainly reduce interference in parameter or gradient space, while replay and distillation often apply the same preservation strength to all previous tasks. These strategies have important limitations. Parameter-space separation does not directly guaranty stable model outputs, average drift can hide severe forgetting on an individual task, and a fixed regularization strength cannot provide different levels of protection for tasks with different degrees of vulnerability. We propose **DualKL**, a task-wise primal–dual approach that preserves previous behavior through predictive-distribution trust regions. After each task, DualKL stores a small set of representative examples together with the model’s output distributions. When learning later tasks, it measures the Kullback–Leibler (KL) divergence between the stored and current distributions for each previous task. Every task is assigned an independent constraint and an adaptive dual variable. The protection strength increases when a task exceeds its allowed drift and remains weak when the task is stable, allowing the model to retain earlier knowledge without unnecessarily restricting new-task learning. We evaluate DualKL on Standard CL, the complete Long Sequence benchmark, and TRACE across multiple language-model architectures. The results show that DualKL consistently improves overall performance and reduces both average and worst-task forgetting compared with replay, parameter-space methods, and fixed functional regularization. These findings demonstrate that selective control in function space is an effective approach for continual language-model adaptation. Our code are available at https://anonymous.4open.science/r/DualKL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.