FBCIR: Diagnosing and Mitigating Partial-Query Reliance in Composed Image Retrieval
Abstract
Composed image retrieval (CIR) is inherently conjunctive: a correct target must satisfy both a reference image and a textual modification. Yet a model can often retrieve the correct target without using the full composed query, because conventional easy negatives can be rejected from only one query component. We refer to this behavior as focus imbalances. We propose FBCIR, a multi-modal focus interpretation method that identifies visual and textual input components sufficient to preserve a model's retrieval decisions. Motivated by this diagnosis, we further develop FBCIR-Data, a CIR data augmentation workflow that enriches existing datasets with curated hard negatives requiring discrimination of both visual and textual constraints, together with a controlled benchmark and finetuning data for evaluating and mitigating focus imbalances. Experiments across representative CIR models reveal substantial vulnerabilities under such hard-negative settings. Under matched finetuning configurations, FBCIR-Data consistently improves hard-case retrieval for VLM-based retrievers and outperforms alternative training datasets, with additional gains observed on standard benchmarks. These results show that explicitly exposing models to negatives that invalidate single-component shortcuts provides an effective way to diagnose and improve robustness in composed image retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.