acceptodds
Under review as a conference paper at ICLR 2027

CrossCheck: Fusing Global Relevance and Fine-Grained Semantic Evidence for Training-Free Composed Image Retrieval

Abstract

Composed image retrieval requires overall target compatibility and precise adherence to a requested modification. However, global multimodal relevance can overlook specific attributes and relations, while fine-grained semantic evaluation may miss broader visual compatibility. We present CrossCheck, a training-free framework that combines these complementary sources of evidence. CrossCheck first estimates global relevance using CLIP-based retrieval and image-text matching. It then constructs semantic checks as target-contrast statement pairs and performs contrastive likelihood scoring: a frozen multimodal large language model compares the conditional token likelihoods of the two statements under the same visual context, yielding fine-grained evidence for individual modification and preservation requirements. We aggregate these signals with sensitivity to weakly supported checks and fuse them with global relevance for final ranking. CrossCheck requires no task-specific training and runs entirely locally with public weights. Across CIRCO, CIRR, and Fashion-IQ and three retrieval backbones, CrossCheck consistently improves over the compared methods. On a fixed CIRCO shortlist, global relevance and fine-grained semantic evidence achieve 29.56 and 30.43 mAP@5 individually, while their fusion reaches 41.60, showing that the gain arises from complementary ranking evidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.