Theoretical Understanding of Dynamic Mixture of Experts in Continual Learning
Abstract
The use of Mixture-of-Experts (MoE) architectures to mitigate catastrophic forgetting in continual learning (CL) has attracted increasing research interest. In particular, dynamic MoE frameworks that progressively expand the expert pool achieve strong empirical performance by allocating dedicated capacity to incoming tasks. Despite this success, a rigorous theoretical understanding of their underlying mechanisms remains largely underexplored. This work provides a formal analysis of dynamic MoE in CL, introducing a framework that integrates expert expansion, freezing, and inheritance. Our analysis reveals that as task diversity grows, expert specialization proceeds in distinct phases, motivating a phased expansion mechanism. We prove that only a finite set of experts is required to cover an infinite task stream within any prescribed error tolerance, justifying an efficient expansion strategy. Furthermore, we derive explicit bounds on forgetting and generalization error, showing that dynamic expansion with expert inheritance ensures stable learning and fast adaptation. Experiments on both synthetic and real-world benchmarks validate our theoretical insights and offer principled guidelines for designing scalable MoE-based CL systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.