DCHA: Dynamic Compositional Adaptation with Hierarchical Drift Constraints for Incremental Vision-Language Object Detection
Abstract
Incremental vision-language object detection (IVLOD) aims to continuously adapt vision-language detectors to emerging domains while preserving previously acquired knowledge and open-vocabulary generalization. While current methods typically optimize LoRA experts to achieve efficient adaptation, these training paradigms exhibit poor plasticity, as empirically evidenced by highly homogeneous parameter update directions across different categories. This rigidity stems from sharing restricted parameter spaces. Specifically, task-level LoRA methods force all categories to share unified experts, while existing category-level LoRA methods rigidly assign experts based solely on labels, completely ignoring intra-class feature variance and imposing a fixed routing capacity. To break this rigidity, we propose category-adaptive composition (CAC) equipped with learnable category-adaptive routing (LCAR). Inspired by Sparsegen, LCAR utilizes learnable keys and thresholds to dynamically select and assemble the most matched fine-grained parameter vectors into exclusive category LoRA experts. This demand-aware routing ensures that update directions are highly tailored to category-specific feature requirements, thereby significantly enhancing parameter plasticity and reducing training interference. Furthermore, because CAC introduces a hierarchical parameter structure where vectors compose category experts and category experts aggregate into task experts, parameter deviations inevitably accumulate across levels. To mitigate this severe error accumulation, we introduce a hierarchical drift constraint (HDC) that performs layer-wise regularization. Extensive experiments on ODinW-13 and ODinW-O demonstrate that our framework consistently outperforms state-of-the-art methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.