CaCLIP: Class-Aware CLIP for Replay-Free Multi-Label Class-Incremental Learning
Abstract
Multi-label class-incremental learning (MLCIL) requires a model to learn new classes over a sequence of tasks while preserving previously acquired knowledge. Adapting CLIP to MLCIL poses two challenges. First, a class-shared visual representation and class-specific text embeddings have mismatched granularities, breaking class-wise visual–textual semantic symmetry. Second, missing negative supervision under task-level partial labels leads to a high false-positive rate (FPR). To address these challenges while mitigating catastrophic forgetting, we propose CaCLIP, a replay-free Class-Aware CLIP framework comprising Semantic Evidence Allocation (SEA) and FPR-Adaptive Loss (FAL). SEA uses text-initialized semantic anchors to decouple shared patch-level evidence into dedicated class-specific visual representations, thereby promoting class-wise visual-textual semantic symmetry. Rethinking recent FPR-suppression strategies from a class-aware perspective, FAL estimates class-specific false-positive risk from confidence and adaptively strengthens negative supervision through its classification and distillation components, reducing FPR across current and old label spaces while mitigating catastrophic forgetting. Extensive experiments on PASCAL VOC, MS-COCO, and the real-world NUS-WIDEseq benchmark demonstrate that CaCLIP achieves state-of-the-art replay-free MLCIL performance with a few trainable parameters and minimal inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.