Which Way Is Broken? Language-Free Zero-Shot Anomaly Detection with Frozen Vision Foundation Models
Abstract
Zero-shot anomaly detection (ZSAD) aims to identify and localize unseen defects without access to reference target samples. Prior approaches rely heavily on Vision-Language Models (VLMs) by matching visual patch tokens with hand-crafted or learnable text prompts. However, this language-grounded setup creates a fundamental bottleneck: visual defects are open-ended and fine-grained, making them difficult to describe accurately in text. Recent language-free methods avoid text, but they change the vision model itself. We show that neither is needed. Frozen vision models naturally separate normal and defective patches along a direction that is shared across datasets. We learn a few such directions for normal and defective patches directly on the frozen features, and score each patch by how much it is closer to the corresponding set. During training, each direction is treated as a small spread of nearby directions rather than a single point. Since defects on unseen products point in slightly different directions, such a direction is more likely to transfer. The method uses no text, changes no model weights, and needs very little training data. Across nine benchmarks in strict zero-shot settings, it matches or exceeds language-based and vision-only methods on AUROC, and improves F1-max by 6.8 points at the pixel level and 1.6 points at the image level. It also performs well in the data-constrained regime, trained on only two images per class, it already matches or outperforms prior methods trained on the full data. It works with six different vision backbones and benefits directly from stronger ones. The source code will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.