DART: Distribution-Guided Adaptive Routing for Federated Continual Learning of Generative Vision-Language Models
Abstract
Federated continual learning (FCIL) enables vision-language models (VLMs) to continually learn from decentralized, evolving data while retaining previously learned knowledge. Although mixture-of-experts (MoE) adaptation supports task specialization, client heterogeneity and sequential updates can bias routing toward dominant or recent tasks, while updates to the shared backbone can degrade historical experts even when they are correctly selected. We propose DART, a Distribution-guided Adaptive RouTing framework that jointly addresses routing bias and backbone-induced forgetting. Clients summarize expert-specific routing inputs in a fixed reference space, while the server maintains per-expert historical distributions and integrates them with current client statistics. Balanced synthetic features sampled from these distributions calibrate a global router without replaying historical raw image and text examples. Meanwhile, historical experts are frozen and backbone updates are constrained to a bootstrapped residual subspace, reducing interference with previously learned representations. Extensive experiments show that DART consistently outperforms state-of-the-art baselines, with performance improvements of up to 7.18%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.