Purifying Generic Visual Scene Graph to Enable Caption-Aligned Image-Text Retrieval
Abstract
Scene graph matching methods for image-text retrieval represent images and captions with structured entities and relations, enabling fine-grained cross-modal matching beyond isolated region-word alignment. However, existing methods typically rely on generic scene graph generators, which produce Visual Scene Graphs (VSGs) containing caption-irrelevant semantics. This creates a semantic gap with the sparse Textual Scene Graphs (TSGs) parsed from captions. To bridge this gap, we propose PSGR, a retrieval-oriented framework that purifies generic visual scene graphs prior to matching. PSGR uses GPT-5-generated caption-guided supervision to adapt a Large Vision-Language Model (LVLM) with SFT and GRPO, then applies Rank-Preserved Candidate Sampling (RPCS) to select informative and diverse triplets from refined and generic graph candidates. The resulting Purified Scene Graphs (PSGs) are compared with TSGs through local matching and global GPO-based graph matching. The LVLM refiner is adapted using Flickr30K training supervision and then frozen to generate offline graph inputs, while the retrieval matchers are trained separately for each benchmark. Under this protocol, PSGR obtains 558.9 and 462.0 Rsum on Flickr30K and MS-COCO 5K, respectively, outperforming the existing scene graph matching methods. With the visual backbone and matching architecture held fixed, LVLM-based graph refinement and RPCS jointly improve Rsum by 13.0 points on Flickr30K and 6.6 points on MS-COCO 5K over generic VSG top-rank sampling. Our curated data and code are available at https://anonymous.4open.science/r/PSGR-DA6D
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.