CLoRE: Drift-Controlled Continual Learning of LLMs via Expert Expansion with Reservoir Routing
Abstract
Continual learning of large language models (LLMs) by mixture-of-experts (MoE) expansion freezes earlier experts and adds new ones, which in principle prevents forgetting. However, the router is shared, and new experts added to it can displace or down-weight an earlier input's experts. Most prior methods constrain routing during or after training. Yet what must be preserved is the representation, which can drift as soon as new experts are added. We propose CLoRE, which controls this representational drift without storing past data. Concurrent router alignment trains the router at every step with the language-modeling loss, matching no previous router. Reservoir routing initializes new experts and their router rows so that expansion preserves the softmax denominator in expectation. Self-generated replay steers each question toward an earlier task by routing only its first token through that task's experts, and replays the backbone. On TRACE with two LLMs (e.g., Llama-3.1 and Qwen3) and on an MoE trained from scratch, CLoRE outperforms every baseline in average accuracy and forgetting, even one using original data, and largely preserves the backbone's general abilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.