Fused Embedding Edit for Zero-Shot Composed Image Retrieval
Abstract
Zero-Shot Composed Image Retrieval (ZS-CIR) retrieves target images using a reference image paired with a modification text, without training on costly real triplet annotations. Existing methods project images into text space or synthesize pseudo-visual features, ultimately performing cross-modal matching that suffers from an inherent *representation gap*, losing fine-grained visual details. We propose **FusionDiff**, a framework that addresses this representation gap through retrieval within a unified joint vision-language embedding space. Its conditional embedding diffusion model learns the distribution of target fusion embeddings conditioned on the composed query. For data efficiency, we design a lightweight Control-Adapter that adapts pre-trained diffusion backbones using only 200K synthetic triplets (*two orders of magnitude less* than prior methods). Experiments on three benchmarks demonstrate state-of-the-art performance: +2.34 mAP@50 on CIRCO, +1.28 R@10 on CIRR, and +1.58 R@50 on FashionIQ over previous best methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.