Imagine-then-Reason for Training-free Zero-shot Composed Image Retrieval
Abstract
Composed image retrieval (CIR) retrieves gallery images given a reference image and a modification instruction, whose joint interpretation is a key challenge. Prior work identifies CIR as a task in which the reference content to preserve, modify, or omit is ambiguous. However, our preliminary analysis of instruction types and retrieval performance highlights an underexplored source of fuzziness: . In particular, comparative instructions such as “make it darker” specify relative changes rather than concrete appearances, making them difficult to translate into self-contained target captions. Caption-based retrieval consistently underperforms on these queries. To address this challenge, we propose Imagine-Then-Reason for Composed Image Retrieval (ITR-CIR), a training-free zero-shot framework that (1) uses a visual target hypothesis to guide reasoning and generate concrete target captions; (2) fuses multimodal target estimators’ evidence for retrieval. ITR-CIR first edits the reference image into a possible target image, then uses an MLLM to assess this hypothesis against the original query and generate grounded target captions while disregarding editing errors. The final ranking combines the captions, the synthetic image, and a direct composition of the original query. Experiments on CIRR, FashionIQ, and CIRCO demonstrate the superior performance of ITR-CIR, achieving the best training-free CIRR and FashionIQ full-gallery recall.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.