AnomalyPrompt: LLM-Guided Adaptive Prompting via Visual Perception for Zero-Shot Anomaly Detection
Abstract
Industrial visual anomaly detection has long been constrained by the scarcity of anomalous samples and the high diversity of defect appearances. Recent CLIP/VLM-based methods have demonstrated promising cross-category generalization for ZSAD through prompt optimization. Most existing methods rely on static or manually crafted prompts or perform only passive adjustments, lacking mechanisms for active prompt exploration and quality assessment. Consequently, they struggle to capture fine-grained distinctions between normal and abnormal states and ultimately fail to consistently generate reliable prompts for anomaly detection. To address this limitation, we propose AnomalyPrompt, a training-efficient vision-language anomaly detection framework that unifies prompt generation, prompt evaluation, and visual-semantic alignment through a multi-stage prompting process. AnomalyPrompt operates through a structured “perception-reasoning-decision-execution” pipeline: it first uses the image-level embedding and patch embeddings of the input image to retrieve global surface features and fine-grained local features from hierarchical concept banks; then, an LLM-based Prompt Generator produces diverse, CLIP-friendly natural-language candidate prompts under structured constraints; next, an LLM Prompt Evaluator scores each candidate in terms of cue faithfulness, CLIP suitability, and clarity/readability, selecting the optimized prompt; finally, the selected prompt is encoded together with learnable normal/abnormal state vectors for anomaly scoring and localization. In this way, AnomalyPrompt enables image-specific prompt adaptation while preserving the semantic transparency of natural-language prompts. Extensive experiments on seven industrial anomaly detection datasets and four medical benchmarks demonstrate that AnomalyPrompt achieves SOTA performance in zero-shot anomaly detection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.