Follow the Edit: Counterfactual Diagnostics and On-Policy Rank Distillation for Composed Image Retrieval
Abstract
Composed image retrieval (CIR) retrieves target images based on the reference image and an editing instruction, satisfying the requested changes while preserving relevant visual content. Conventional retrieval metrics measure whether targets appear among the top-ranked results but do not directly assess responses to changes in the user's editing instruction. We introduce CIR-QDiag, a dataset of 10,131 local instruction edits, together with two complementary metrics, Directional Response Rate (DRR) and Calibrated Response Rate (CRR), to assess instruction following in CIR through models’ responses to these edits. Our diagnostic reveals that high retrieval recall does not guarantee high scores on both metrics and that instruction sensitivity varies across models. To improve CIR performance by transferring instruction-sensitive relevance judgments from stronger models to smaller retrieval models, we propose On-Policy Rank Distillation (OPRD), which matches a frozen teacher's soft relevance distribution over student-mined candidates. The student minimizes reverse KL on the shared candidates, focusing supervision on its current rankings. This transfers soft ranking information beyond the annotated target, while retaining student-only inference. OPRD trains on original CIR pairs without diagnostic edits. With BGE-MLLM-S2 as the teacher, OPRD achieves higher CRR than the InfoNCE baseline across all four student checkpoints. For CLIP-L, it improves CRR from 2.00% to 6.22% and CIRR Recall@1 from 12.80% to 42.72% relative to the pretrained checkpoint, with additional retrieval gains on PinPoint and FashionIQ. Code: https://anonymous.4open.science/r/QDiag-OPRD-877B/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.