acceptodds
Under review as a conference paper at ICLR 2027

UniCIR: Unlocking Universal Embeddings for Training-Free Composed Image Retrieval

Abstract

Composed image retrieval (CIR) searches for images that satisfy a textual modification while preserving relevant content from a reference image. Many training-free methods construct queries through generated descriptions, making retrieval depend on the visual details captured in text. Universal multimodal embeddings jointly encode the reference image and modification, providing a strong alternative. However, retrieval can remain visually close to the reference while missing a requested change. We propose UniCIR, a training-free method that refines this joint representation through a controlled description pair. A frozen vision-language generator describes the intended target. We then replace one edit-linked value to form a negative description, keeping all other wording fixed. The shared retriever encodes both descriptions and refines the original query by adding the weighted positive and subtracting the weighted negative. This contrast targets the requested distinction while preserving the original joint embedding as an anchor. The resulting vector searches a fixed gallery with all model weights frozen. Our experiments show gains of 6.31, 7.89 and 8.68 percentage points over Qwen-8B on CIRR test R@1, CIRCO test mAP@5 and standard FashionIQ validation average R@10, respectively. Gains extend to four additional retrievers, while matched ablations show improvements over augmentation with positive descriptions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.