acceptodds
Under review as a conference paper at ICLR 2027

What Does Continual Retrieval Retain? A Candidate-Provenance Study

Abstract

Continual vision-language learning aims to acquire new retrieval capabilities while preserving those learned earlier. Yet evaluating each task within its own gallery leaves an important question open: do conclusions drawn from task-local evaluation still hold when candidates from multiple domains compete in a shared gallery? We evaluate CLIP ViT-B/32 under sequential fine-tuning and Mod-X across five domains at pre-task, post-task, and final checkpoints, using fixed task-local (Local) and five-domain (Shared) galleries. On locally assembled splits, queries, benchmark-defined positives, and candidate identities remain fixed across checkpoints. Across both configurations and retrieval directions, the same later updates produce 0.51–1.14 percentage points more R@10 loss in Shared than Local, averaged equally over the four tasks with subsequent training. We also follow queries successful at R@10 in both galleries immediately after a task. Among their final Shared failures, approximately 29–32% still succeed in Local when within-domain proportions are averaged equally for each configuration and direction. Under a fixed budget of 512 added negatives, the expected R@10 loss differs by source and can increase or decrease after later updates. These findings show that task-local evaluation alone can miss losses of previously successful retrievals under shared candidate competition. Assessing what continual retrieval retains therefore requires considering the candidate environment: how much post-task success survives subsequent updates, and which queries remain successful when domains share a search space.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.