Sharpen: Bringing the Requested Change into Focus for Composed Image Retrieval
Abstract
Composed image retrieval searches for an image that preserves relevant content from a reference image while showing a change specified by a modification text. In the training-free describe-then-retrieve paradigm, a multimodal large language model (MLLM) writes a target description, and Contrastive Language-Image Pretraining (CLIP) ranks database images by similarity to it. The description can favor unchanged context, placing reference-like images that miss the change above the target. A description-based similarity score also cannot compare a candidate directly with the reference to check what it preserves. To bring the requested change into focus, we propose Sharpen, which adjusts scores across the database and within the selected candidate set. Across the database, the global adjustment strengthens the requested change in the query embedding and suppresses popular images that resemble many composed queries before selecting candidates. Within the candidate set, the local adjustment reuses this MLLM to judge the change without the reference, and the complete request with it. Both the change judgment and the reference-conditioned judgment contribute to the final order. With a 4B open-weight MLLM and no training, Sharpen achieves state-of-the-art training-free results on three widely used benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.