acceptodds
Under review as a conference paper at ICLR 2027

D²C-MCL: Dual Decoupling and Consolidation for Multimodal Continual Learning in Vision-Language Models

Abstract

Continual learning for Vision-Language Models (VLMs) aims to acquire new downstream knowledge while preserving previously learned knowledge and pre-trained zero-shot capabilities. However, existing methods largely overlook the evolving roles of different modalities during continual learning. We observe that the relative contributions of visual and textual modalities are inherently dynamic and task-dependent, varying as new tasks are introduced and knowledge accumulates. Motivated by this observation, we propose D²C-MMCL(Dual Decoupling and Consolidation for Multimodal Continual Learning), a dual-decoupling and knowledge-consolidation framework that disentangles continual knowledge along modality and task dimensions. By dynamically modeling modality contributions, D²C-MMCL performs modality-aware cross-task consolidation and selectively exploits modality-specific knowledge for different tasks.Extensive experiments across various continual learning settings demonstrate that our method consistently outperforms strong existing approaches, achieving superior continual learning performance while better preserving zero-shot generalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.