acceptodds
Under review as a conference paper at ICLR 2027

0-CLIP: Training-free Anomaly Detection Hidden in CLIP's Projection

Abstract

CLIP-based zero-shot anomaly detection commonly improves dense discriminability through prompts or adapters trained on auxiliary data. However, reliable transfer to unseen domains remains a practical challenge. This motivates a closer examination of how much anomaly information is already present in CLIP itself. Our preliminary experiments show that CLIP's frozen visual projector contains useful dense anomaly structure, but its patch-level scores tend to rank normal and anomalous regions in the wrong order. Based on this observation, we introduce 0-CLIP, a training-free method that converts into a patch-to-text anomaly operator. The construction reverses the normal–abnormal score polarity, forms a compact operator from the principal singular directions of , and selectively recovers lower-energy directions using a separation criterion computed on the first unlabeled test image of each category. Without any labeled data, the resulting operator yields both pixel-level anomaly maps and image-level anomaly scores. Across industrial and medical benchmarks, 0-CLIP achieves performance competitive with training-based methods, and ablation studies show the effectiveness of the proposed method. These results suggest that substantial dense anomaly discriminability is already latent in CLIP's frozen visual projector and can be exposed without training. Our code is available in the Supplementary Material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.