Reset or Reuse? Understanding How Optimizer Memory Shapes Continual Mixture-of-Experts Learning
Abstract
Continual fine-tuning of mixture-of-experts (MoE) models, a paradigm that sequentially adapts a pre-trained MoE model to a stream of tasks, has become a popular method in lifelong model deployment because it acquires new capabilities without retraining from scratch. However, standard sequential fine-tuning carries optimizer state across task boundaries and updates the full parameter space, so cross-task interference from inherited optimization history and from heterogeneous routed experts and always-active modules is neither measured nor controlled, while existing methods instead estimate parameter importance at substantial storage cost. To address this, we investigate the plasticity–forgetting dynamics of continual MoE fine-tuning through boundary-state interventions and role-structured updates, finding that interference degrades both what a new task acquires and what earlier tasks retain, which motivates separating what to discard from what to keep. Accordingly, we propose MoRE, a structured optimizer-update scheme for moment reset and reuse that clears the optimizer's fast state at task boundaries, reuses accumulated moment history to adapt the effective learning rate, and allocates subsequent updates across granularities, from role-level gates to per-coordinate modulation. On a six-task sequence from the TRACE continual learning benchmark, MoRE averages on OLMoE-Base, versus for sequential fine-tuning (SeqFT) and for the parameter-regularization baseline EWC; on Qwen, it reaches , versus for SeqFT and for EWC; it requires no Fisher estimation pass, and reduces storage for retained coordinate history by relative to EWC's Fisher-plus-anchor state at equal coverage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.