PhraseBind: Binding Phrases to Visual Evidence for Frozen Vision-Language Models
Abstract
Global image–text similarity can overlook which objects carry an attribute or par- ticipate in a relation. We introduce PhraseBind, a phrase-evidence branch that augments a frozen vision-language dual encoder. It decomposes captions into object, attribute–object, and relation phrases, routes each phrase to sparse image patches, and combines confidence-weighted local support with the unchanged global score. Counterfactual captions train binding discrimination, while phrase- region annotations supervise the supporting evidence maps. Under matched training and readout settings, PhraseBind improves ARO-VG accuracy over token- and phrase-local controls by 10.64 and 9.38 points, respectively, while retaining 99.46% of frozen-correct retrieval decisions. The learned evidence also transfers to Ref- COCOg, reaching 48.5% test pointing accuracy without target-dataset supervision, compared with 34.3% for a frozen token–patch scorer. Experiments on CLIP, Open- CLIP, and SigLIP show consistent grounding gains with nearly unchanged retrieval. The same interface also complements reproduced hard-negative objectives, adding spatial evidence to their image–text matching pipelines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.