Rex-MoE: Knowledge Reuse in Expandable Mixture-of-Experts for Continual Learning of Large Language Models
Abstract
Continual learning (CL) requires models to learn new tasks sequentially while preserving previously acquired knowledge. Expandable mixture-of-experts (MoE) provides a natural framework by introducing new experts for incoming tasks while preserving historical experts to retain knowledge. However, existing methods typically reuse historical experts only as whole modules, which limits the flexibility to reuse and recombine useful components of acquired knowledge when learning new tasks. By reusing these components while keeping historical experts fixed, models could better balance new-task learning and knowledge retention. In this paper, we propose knowledge Reuse in expandable MoE (Rex-MoE), which enables flexible knowledge reuse during expert learning while coordinating the contributions of new and historical experts. For expert learning, Rex-MoE factorizes each adaptation into expandable shared bases and a task-specific transformation coupling historical and new bases. This structure allows new tasks to reuse historical bases without rewriting the transformations among them. For expert integration, Rex-MoE learns the relative contributions of new and historical experts while constraining router updates to limit interference. Extensive experiments on CL benchmarks demonstrate state-of-the-art performance of Rex-MoE.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.