acceptodds
Under review as a conference paper at ICLR 2027

PhraseBind: Binding Phrases to Visual Evidence for Frozen Vision-Language Models

Abstract

Global image–text similarity can overlook which objects carry an attribute or par- ticipate in a relation. We introduce PhraseBind, a phrase-evidence branch that augments a frozen vision-language dual encoder. It decomposes captions into object, attribute–object, and relation phrases, routes each phrase to sparse image patches, and combines confidence-weighted local support with the unchanged global score. Counterfactual captions train binding discrimination, while phrase- region annotations supervise the supporting evidence maps. Under matched training and readout settings, PhraseBind improves ARO-VG accuracy over token- and phrase-local controls by 10.64 and 9.38 points, respectively, while retaining 99.46% of frozen-correct retrieval decisions. The learned evidence also transfers to Ref- COCOg, reaching 48.5% test pointing accuracy without target-dataset supervision, compared with 34.3% for a frozen token–patch scorer. Experiments on CLIP, Open- CLIP, and SigLIP show consistent grounding gains with nearly unchanged retrieval. The same interface also complements reproduced hard-negative objectives, adding spatial evidence to their image–text matching pipelines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.