Low Rank, Lasting Knowledge: Mitigating Catastrophic Forgetting in LoRA Fine-Tuning
Abstract
Low-Rank Adaptation (LoRA) adapts large language models to a target domain (TD) with few trainable parameters, but applying the adapter to every input can degrade performance on out-of-target-domain (OTD) data. The base model itself remains unchanged. We address this non-target interference through DG-LoRA, a sample-level dynamic gating method for the post-hoc conditional activation of a single frozen LoRA adapter. A lightweight multilayer perceptron computes a scalar gate value from an answer-independent routing sequence and uses it to interpolate the output logits of the base and LoRA-enabled models before softmax. This value stays fixed throughout candidate scoring and generation. After training the adapter, we freeze both predictive branches and train only the gate, combining domain classification, target-task and non-target preservation losses. These objectives account for domain membership as well as the effect of adaptation on predictions. Across the evaluated models and tasks, DG-LoRA performed comparably to standard LoRA on TD data while substantially reducing non-target interference on OTD data, including general-domain and other domain-specific tasks. OTD performance generally returned close to the base-model level. The gate added fewer than as many trainable parameters as the LoRA adapter.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.