acceptodds
Under review as a conference paper at ICLR 2027

Beyond Representation Drift: Decomposed Relative Geometry Preservation for Continual Vision-Language Learning

Abstract

Continual adaptation of contrastive vision-language models leads to catastrophic forgetting, accompanied by substantial drift in their visual and textual representations. Existing methods mitigate forgetting by constraining changes in model parameters, feature representations, or cross-modal affinities, yet they do not explicitly separate changes that preserve the joint image–text geometry from those that alter it. We find that this distinction is substantial in practice: across a seven-task continual sequence, nearly half of the accumulated representation drift is attributable to a shared orthogonal transformation that leaves image–text retrieval unchanged. Motivated by this observation, we propose Decomposed Relative Geometry Preservation (Decomp-RGP), which measures drift after factoring out shared orthogonal transformation through closed-form Procrustes alignment and exactly decomposes the remaining residual into visual deformation, textual deformation, and cross-modal orthogonal alignment mismatch. We further introduce a budget-based adaptive weighting scheme to control the contributions of these components in fine-tuning process. Across seven image–text retrieval tasks, Decomp-RGP improves Mean R@1 by 5.32% over the strongest baseline, while maintaining stronger zero-shot performance across multiple CLIP backbones.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.