acceptodds
Under review as a conference paper at ICLR 2027

Learn the (Register) Gap: Making Effective Specialist Text-to-Image Retrieval with Generalist Encoders

Abstract

Frozen vision-language encoders often underperform on text-to-image retrieval in specialist collections, even in domains well represented in web-scale pretraining. We show that two small linear maps, one for captions and one for images, trained on the collection's own caption-image pairs with the encoder frozen, recover much of this loss. The maps improve recall in 57 of 58 encoder-collection pairs spanning medicine, remote sensing, biodiversity, art, science, fashion and news, by up to 34 points at R@10, and more than doubling it in eight of them. The frozen encoder's own ranking of the catalog can predict the gain before training, and maps learned on one collection transfer only to collections that share its conventions of description and depiction. The failure is therefore a register gap, not a domain gap: the specialist features are present, but specialist ways of describing and depicting leave captions and images misaligned. The maps recover half or more of LoRA's gain at a fraction of its cost, and add to low-rank LoRA when applied on top. Applied to a caption's phrases and an image's patches, the same maps improve retrieval further and ground phrases in images without box annotations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.