acceptodds
Under review as a conference paper at ICLR 2027

CADT: Category-aware Adaptation with Disentangled Transformations for Incremental Vision-Language Object Detection

Abstract

Incremental vision-language object detection (IVLOD) enables detectors to continuously adapt to new downstream domains while preserving pre-trained open-vocabulary generalization. However, existing methods still suffer from parameter interference during continual training. Specifically, this interference manifests in two dimensions. Regarding inter-category interference, existing category-level adaptation methods typically learn update matrices in a single parameter space, coupling the update direction and amplitude. From an optimization perspective, this coupling mechanism severely limits the degrees of freedom for category-specific adjustments. Consequently, dominant gradients from shared generic features easily overwhelm the signals of category-unique features, forcing category experts to update along generic features and resulting in severe homogenization. Regarding intra-category interference, recurring categories across tasks are updated with different task contexts, causing representation drift to accumulate task-by-task, which exacerbates forgetting. To address these issues, we propose a category-aware adaptation framework with disentangled transformations (CADT). For inter-category interference, we propose disentangled category-aware transformation (DCT). Utilizing the structure of SVD, DCT explicitly decouples the update direction and amplitude for separate optimization. This mechanism releases the degrees of freedom for parameter updates, breaks the homogenization bottleneck of optimization, and enables the independent adjustment of adaptation direction and intensity to learn category-specific features. For intra-category interference, we design category adaptation consistency (CAC), which fuses the historical and current category parameter and imposes consistency constraints to avoid excessive parameter deviation. Experiments on ODinW-13 and ODinW-O show that CADT achieves state-of-the-art performance, striking a favorable balance between category-specific adaptation and knowledge preservation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.