Beyond Cosine Similarity: Uncovering Anomaly Information in CLIP's Raw Embedding Geometry
Abstract
Extensive experiments following the image-text cosine similarity pre-training paradigm of Contrastive Language-Image Pre-training (CLIP) indicate that CLIP excels at image classification rather than fine-grained anomaly detection; consequently, a prevailing consensus holds that CLIP is anomaly-agnostic. Existing studies have concentrated on downstream training to better adapt CLIP for anomaly detection tasks; in other words, image-text cosine similarity is the computational foundation underlying CLIP-based anomaly detection. In this paper, we investigate the directional information about anomalies in CLIP's raw embedding space and the limitations of cosine similarity with the idea of computing geodesics distance for anomaly shifts. We investigate the raw representation of CLIP and find that anomaly information is already encapsulated in CLIP's raw embedding space, which is mainly formed by class-dependent directional variation. The key reason why CLIP cannot recognize the normality/abnormality in images is that the now-ubiquitous cosine similarity ignores the anisotropy of the normal/abnormal distribution. We demonstrate this through controlled geometric analysis, and use a top-down statistical method as the simplest parameter-free update to uncover the latent anomaly information, making this process interpretable. In extensive experiments, CLIP's latent anomaly information substantially improves AD performance over image-text cosine similarity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.