LoCaR: Localized Canvas Retrieval With Frozen Vision-Language Models
Abstract
Whole-image embeddings can under represent small objects and local attributes that distinguish relevant images for localized queries. We present LoCaR, a training- free framework for localized text-to-image retrieval with frozen vision–language models. LoCaR constructs and utilizes independent proposals (SAM2) focusing on the object so as to get a more localized representation that independently encodes each region using the frozen visual encoder. We further study paired-support adaptation as an optional residual that uses matched text–region exemplars without updating the underlying encoders. We also introduce LocFG, a Visual Genome- derived benchmark designed to evaluate localized retrieval for smaller and more challenging referents. Across several vision language backbones, our results show that improving the region representation itself is the main source of gain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.