acceptodds
Under review as a conference paper at ICLR 2027

Diffusion Prior-driven Pseudo Target Extraction for Training-Free Zero-Shot Composed Image Retrieval

Abstract

Zero-shot Composed Image Retrieval (ZS-CIR) aims to retrieve relevant images from a gallery using both a reference image and manipulation text without requiring task-specific training. Recent approaches imagine the target by generating a pseudo target image in pixel space with an image generation model such as Qwen-Image-Edit, and then extract its embedding with the CLIP image encoder. However, the pseudo target image may introduce redundant image details that mislead retrieval, and the image generation and re-encoding incur substantial computational overhead. To address these issues, we propose Diffusion Prior-driven Pseudo Target Extraction (PPT), a training-free framework that extracts the pseudo target embedding directly with a diffusion prior in the CLIP image embedding space. Without any diffusion decoder or CLIP image encoder at query time, PPT achieves competitive performance on object/scene manipulation and attribute manipulation tasks with much lower latency. Our code will be publicly released upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.