acceptodds
Under review as a conference paper at ICLR 2027

How Much to Expand, and How to Route: Dynamic-Capacity Mixture-of-Experts for Continual Learning

Abstract

Continual learning (CL) adapts large pre-trained models to a stream of new tasks without revisiting earlier data. A common architecture adds a group of LoRA experts per task, freezes old experts, and routes each token through a learned Mixture-of-Experts (MoE) router. Freezing prevents weight overwriting, but forgetting persists through routing drift as old-task tokens shift toward newer experts. Two decisions govern this framework: how much capacity each task adds, and how the accumulated experts are routed at inference. Existing methods fix the first by hand on both the expert-count and rank axes, and address the second with per-token routing alone or with trained task predictors. We propose DyCapMoE, which makes both decisions data-driven. Dynamic Expert Capacity (ECap) attaches a Hard-Concrete gate to each new expert with a sign-aware coupling that keeps gate closure monotone under softmax routing, and Dynamic Rank Capacity (RCap) applies the same mechanism to each expert's rank. Gates that reach zero are pruned at task end, so each task retains only the capacity it uses. A prototype task router provides the task-level inference prior from by-products of the forward pass without trainable parameters. Across four CL benchmarks on LVLMs and LLMs, the learned budget retains under a third of the allocated rank at equal or better final accuracy, the task router brings forgetting to near zero, and the full model attains the best final accuracy on all four benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.