acceptodds
Under review as a conference paper at ICLR 2027

Imagine-then-Reason for Training-free Zero-shot Composed Image Retrieval

Abstract

Composed image retrieval (CIR) retrieves gallery images given a reference image and a modification instruction, whose joint interpretation is a key challenge. Prior work identifies CIR as a task in which the reference content to preserve, modify, or omit is ambiguous. However, our preliminary analysis of instruction types and retrieval performance highlights an underexplored source of fuzziness: . In particular, comparative instructions such as “make it darker” specify relative changes rather than concrete appearances, making them difficult to translate into self-contained target captions. Caption-based retrieval consistently underperforms on these queries. To address this challenge, we propose Imagine-Then-Reason for Composed Image Retrieval (ITR-CIR), a training-free zero-shot framework that (1) uses a visual target hypothesis to guide reasoning and generate concrete target captions; (2) fuses multimodal target estimators’ evidence for retrieval. ITR-CIR first edits the reference image into a possible target image, then uses an MLLM to assess this hypothesis against the original query and generate grounded target captions while disregarding editing errors. The final ranking combines the captions, the synthetic image, and a direct composition of the original query. Experiments on CIRR, FashionIQ, and CIRCO demonstrate the superior performance of ITR-CIR, achieving the best training-free CIRR and FashionIQ full-gallery recall.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.