acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Zero-Shot Capabilities in Vision Foundation Models

Abstract

Frozen vision foundation models are used zero-shot on tasks they were never trained for, such as matching features across images for semantic and geometric correspondence, tracking points through video, and nearest-neighbour classification. Evaluations extract features, compare them by cosine similarity, and choose the candidate with the highest score. Cosine similarity does not account for how variance is distributed across feature directions. In many models, a few directions carry most of the variance. We measure this concentration by the effective dimension of the covariance of the extracted feature vectors. Across 36 models and 46 extraction points, ranges from to , despite feature dimensions ranging from to . Such low effective dimensions matter because the cosine similarity between unrelated features has variance of about , so when is small, wrong candidates can score almost as highly as the correct match. We therefore whiten the features with their mean and covariance, estimated on unlabeled images, before computing cosine similarity. Whitening requires no labels or training and has long been standard in image retrieval, yet it is not part of the usual zero-shot evaluation of foundation-model features. Across eleven dense benchmarks, whitening improves performance on of model–benchmark pairs, by points on average and up to PCK points. The gain is larger for models with lower pre-whitening . Diffusion transformers gain the most. Their raw features match poorly under cosine similarity, but after whitening they achieve matching performance comparable to strong self-supervised encoders. Whitening should be part of the zero-shot evaluation of frozen vision features.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.