Understanding Candidate-Shared Evidence in Visual Foundation Models
Abstract
What parts of a CLIP feature matter for a prediction? We find that once the classifier has narrowed an image to a few likely classes, most of the feature can be removed without changing the answer. We take the top- classes from a frozen classifier and use a sparse autoencoder to split the CLIP feature into concepts. Across 11 datasets, the concepts that all candidates share are also the most strongly expressed, and they make up roughly two thirds of the feature. We show why the classifier trained on top of CLIP keeps them: its cross-entropy loss over all classes rewards evidence that separates the likely classes from all the others, even when that evidence does nothing to tell the likely classes apart. We then train two small rerankers that only choose among the candidates, one on the full feature and one with the shared concepts deleted. Both improve three standard few-shot CLIP classifiers, a linear probe and two prompt-learning methods, by 1.3 to 5.9 points at 16 shots, and they match within 0.2 points. The deleted concepts alone reach 13% accuracy among the candidates against 68% for what remains, and the full-feature reranker learns to ignore them on its own, cutting their influence on the score by 74%. Our results show that global recognition relies on evidence the candidates share, and that local discrimination can discard it and learns to do so on its own.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.