acceptodds
Under review as a conference paper at ICLR 2027

Don't Care About Encoders: Encoder-Agnostic Class-Incremental Cross-Modal Retrieval

Abstract

Real-world retrieval systems must continuously evolve to incorporate new semantic classes and upgraded backbone architectures while maintaining retrieval across modalities. However, existing class-incremental learning (CIL) methods often assume fixed encoder architectures, while backward-compatible training typically focuses on unimodal scenarios. To address this challenge, we propose Don't Care About Encoders (DCAE), an encoder-agnostic framework designed for class-incremental cross-modal retrieval. DCAE uses lightweight residual adapters on top of frozen backbone encoders to map feature vectors into a shared normalized representation space. To promote geometric stability across modalities and encoder versions, we anchor representations to fixed class prototypes with simplex equiangular tight frame (ETF) geometry. Our method further encourages preservation of backbone relationships through a relational matching loss and mitigates catastrophic forgetting via supervised contrastive distillation and replay. Experimental results on audio-visual retrieval tasks demonstrate that DCAE outperforms the evaluated baselines in final retrieval accuracy, while compatibility matrices assess retrieval against old gallery embeddings without re-encoding the corresponding items. The code will be made available after acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.