acceptodds
Under review as a conference paper at ICLR 2027

Zero-Drift Expert Pool: Factorized Routing for Stable Pre-Trained-Model Continual Learning

Abstract

Continual learning with pre-trained models (PTMs) keeps the backbone frozen, adds lightweight adapters, and stores no exemplars. A standard way to scale such a learner is to grow a pool of experts, one per task, and route inputs among them. The obstacle is routing stability: any scheme that normalizes scores over the expert set (softmax, sparsemax, normalized sigmoid) changes every old gate when a new expert is appended, re-mixing previously frozen features. We prove that any normalized router drifts, that this drift cannot be eliminated via regularization, and that factorized per-expert gates eliminate gate-level drift. Building on this, we introduce Zero-Drift Expert Pool (ZDEP): factorized sigmoid gates with hard top-1 routing so appended experts do not affect existing gates, a validation-based controller that adds or reuses an expert, and a modular classifier decoupled from the router. On CIFAR-100, ImageNet-100 (10 tasks each), and ImageNet-R (5 tasks) with an ImageNet-21k ViT-B/16 backbone, ZDEP reaches 0.902, 0.935, and 0.782 average accuracy, surpassing RanPAC (0.874/0.892/0.660), FeCAM (0.851/0.901/0.381), SimpleCIL, and CL-LoRA. A learned softmax router (SoftMoE-style) achieves only 0.358 on CIFAR-100; by contrast ZDEP holds routing drift at , a nine-order-of-magnitude difference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.