Gated Modality Fusion for CLIP-based Class-Incremental Learning
Abstract
Class-Incremental Learning (CIL) enables models to continuously learn new classes without forgetting previously acquired knowledge. Due to the rich and transferable visual–text representations learned by Contrastive Language–Image Pre-training (CLIP) from large-scale image–text data, CIL methods employing CLIP as a pretrained backbone have recently achieved impressive performance. However, visual and textual representations provide complementary discriminative cues, whose relative effectiveness varies across class pairs. Most existing methods typically combine them using a fixed global ratio, implicitly assuming that their relative reliability is invariant across inputs and classes. We empirically show that this assumption is suboptimal: the optimal fusion ratio varies across class pairs and may change as the representation space evolves during incremental learning. In this paper, we propose GMF-CLIP, a Gated Modality Fusion framework that adaptively integrates CLIP’s visual prototypes and textual classifiers through a learnable gating mechanism. For each input, GMF-CLIP dynamically estimates the modality contribution for each candidate class, enabling sample- and class-conditioned fusion. To sustain reliable gated fusion as the visual representation space evolves, we use the Visual Prototype Refresh (VPR) mechanism tailored to the gating network, which leverages changes in the learned fusion ratios to estimate representation shifts and calibrate old-class prototypes. Together, the gated fusion and prototype refresh mechanisms improve classifier reliability and reduce interference between old and new classes. Experimental results on four commonly used datasets demonstrate that GMF-CLIP achieves state-of-the-art performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.