acceptodds
Under review as a conference paper at ICLR 2027

When Do Local Features Help? Patch-Area Matching in Fine-Grained Vision–Language Models

Abstract

Contrastive vision–language models (VLMs) align a single global image embedding with text, capturing whole-image structure but discarding local detail that fine-grained recognition depends on. Frozen patch features retain localizable structure that text-based scoring misses, while gradient-based adaptation risks eroding transferable features. We propose GL-VLM, a training-free, few-shot method that keeps the backbone frozen and augments the global embedding with a local "patch-area" descriptor: per class, we cluster the support patches into region prototypes and score a query by how its patches fall into each class's areas, combined with a global prototype. Stacked on a strong Tip-Adapter, it improves texture/part/scene categories (mean +1.4 points over 15 backbones, up to +5.4 at k=16) and never hurts (a β=0 fallback). The benefit is interpretable and stated a priori: local patch-areas help where distinctions are textural, part-based, or scene-level, and are neutral where text is near-perfect or the cue is global shape text priors and local structure are complementary, and which wins is predictable from the category. Because the backbone is never updated, GL-VLM preserves zero-shot transfer exactly, at zero training cost. We do not claim state of the art covariance-based methods match or exceed it but characterize when local features help.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.