Diffusion Prior-driven Pseudo Target Extraction for Training-Free Zero-Shot Composed Image Retrieval
Abstract
Zero-shot Composed Image Retrieval (ZS-CIR) aims to retrieve relevant images from a gallery using both a reference image and manipulation text without requiring task-specific training. Recent approaches imagine the target by generating a pseudo target image in pixel space with an image generation model such as Qwen-Image-Edit, and then extract its embedding with the CLIP image encoder. However, the pseudo target image may introduce redundant image details that mislead retrieval, and the image generation and re-encoding incur substantial computational overhead. To address these issues, we propose Diffusion Prior-driven Pseudo Target Extraction (PPT), a training-free framework that extracts the pseudo target embedding directly with a diffusion prior in the CLIP image embedding space. Without any diffusion decoder or CLIP image encoder at query time, PPT achieves competitive performance on object/scene manipulation and attribute manipulation tasks with much lower latency. Our code will be publicly released upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.