acceptodds
Under review as a conference paper at ICLR 2027

Revisiting Continual Learning for Multimodal Retrieval

Abstract

Retrieval is a core capability of vision-language models (VLMs), yet continual adaptation for retrieval remains underexplored, with existing evaluations largely restricted to narrow domains and a single training protocol. We introduce SCOPE, a reproducible framework for continual multimodal retrieval spanning seven heterogeneous visual domains, three training protocols, and both task-specific and global-gallery evaluation. Using SCOPE, we find that conclusions about continual learning (CL) methods depend strongly on the evaluation setting: method rankings change across training protocols, task-specific galleries conceal cross-task retrieval failures, and simple fine-tuning remains competitive with dedicated CL methods. Our analysis further shows that preserving local neighbourhood structure is insufficient to preserve fine-grained retrieval rankings. As a reference method for SCOPE, we introduce Adapter Routing and Consolidation (ARC), which limits cross-task interference through low-rank adapters and selective consolidation. ARC consistently outperforms the evaluated methods across SCOPE settings. Together, these results establish SCOPE as a comprehensive evaluation framework and highlight the importance of diverse training protocols and gallery settings for reliably assessing progress in continual multimodal retrieval.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.