Pointwise Vision-Language Reranking with Reinforcement Fine-tuning for Composed Image Retrieval
Abstract
Composed Image Retrieval (CIR) requires retrieving a target image given a reference image and a modification text describing desired attribute changes. Existing approaches predominantly rely on embedding-based methods that fuse visual and textual features into a joint representation for retrieval. However, these methods struggle with fine-grained semantic understanding of the modification, especially when the desired change is subtle or compositional. We propose a pointwise vision-language model (VLM) reranker trained with reinforcement fine-tuning that operates on top of an off-the-shelf coarse retriever. Specifically, our method frames multi-candidate ranking as a policy optimization problem: the VLM assigns a relevance score to each candidate image through next-token prediction of "Yes/No" tokens, and we train the model via a GRPO-style objective whose stage0-aware reward is computed over rankings that fuse the learned scores with the coarse retriever's prior. Additionally, at inference the learned pointwise scores are fused with the stage0 ranking through a rank-log transformation, enabling the model to both exploit its own semantic understanding and preserve the retriever's recall. Because the protocol consumes only the retriever's top- candidate lists, the same recipe attaches to different coarse retrievers without re-encoding their galleries; we instantiate it with DQU-CIR and MCoT-MVS and train a separate adapter for each. On FashionIQ, Shoes, and CIRR, three benchmark datasets for composed image retrieval, PVR-RFT consistently improves the stage0 retriever it is paired with, raising the controlled top-256 R@10 on all FashionIQ categories and the official CIRR test-server metrics over their matched stage0 baselines, and comprehensive ablation experiments validate the contribution of each component. Our source code and complete training/evaluation configuration are available in an anonymous repository at https://anonymous.4open.science/r/PVRRFT-code.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.