acceptodds
Under review as a conference paper at ICLR 2027

The Class Prior as a Domain-Dependent Bias: Layer-Wise Visual Prompt Synthesis for Frozen-CLIP Zero-Shot Anomaly Detection

Abstract

CLIP-based zero-shot anomaly detection treats the class-semantic prior as a resource the text branch must supply, whether through class names, handcrafted templates, or per-class descriptions. This paper argues that under cross-domain transfer the prior is instead a domain-dependent bias term. Replacing the class-name slot with a per-layer, image-conditioned prototype read-out stays within 0.5 points of the strongest conditional-prompt baseline, while the class-name prompt itself trails that read-out by 3.7 points, so the prior is not a requirement. Its removal yields 5.1/6.6 points at pixel level in the medical domain against 0.3/0.1 in the industrial domain, an order-of-magnitude difference between two domains evaluated under one protocol, which is the pattern predicted by a bias term and not by a method that is uniformly stronger. Within the prompt the learnable component is what carries the adaptation: discarding it costs 8.8/17.0 points, so what proves indispensable is the prompt's optimizable position in CLIP text space rather than what it says. Because each tapped visual layer synthesizes its own prompt, the construction is structurally distinct from prior methods, which match multi-layer evidence against a single prompt. On 14 industrial and medical benchmarks under a unified protocol, the method matches or exceeds the strongest baselines in accuracy while using 1284 GFLOPs against AdaCLIP's 5570, a factor of 4.3, and running at 12.0 FPS against 5.5, with neither a class name nor a handcrafted state word at inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.