Just Enough Experts for Continual Learning of Large Language Models
Abstract
Recent studies on continual learning (CL) have shown that data replay is a simple yet effective strategy for mitigating catastrophic forgetting in large language models (LLMs). However, its effectiveness depends heavily on the quantity and quality of rehearsal data, and substantial forgetting can persist when either is inadequate. Moreover, updating shared model parameters across tasks can introduce interference and further degrade performance. To address these limitations, we propose Adaptive Expansion and Identification (AEI), a framework that combines minimal yet sufficient expert expansion with adaptive expert identification. Specifically, AEI encourages newly introduced experts to capture feature directions complementary to those represented by existing experts. As the represented feature subspace grows across tasks, the remaining complementary space becomes increasingly constrained, encouraging compact and effective expert expansion with reduced redundancy and cross-task interference. Meanwhile, AEI employs a context-aware gating module trained with task-identity supervision to adaptively identify and combine the independent experts relevant to each input. Together, these mechanisms support efficient adaptation to new tasks while preserving previously acquired knowledge. Extensive experiments across multiple LLM backbones demonstrate that AEI outperforms competing continual learning methods using only 1% rehearsal data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.