CRAG-CLIP: Cross-Layer Relational Abnormality Geometry for Zero-Shot Image Anomaly Detection
Abstract
CLIP-based zero-shot image anomaly detection has achieved encouraging progress by leveraging transferable vision–language representations to identify anomalies in previously unseen datasets. Recent source-only approaches further improve cross-dataset transfer by learning anomaly-aware text prompts from auxiliary source data. However, existing prompt-based approaches typically represent the abnormal concept with a single normal-abnormal opposition and fuse multi-layer visual evidence through fixed heuristics. This paradigm inevitably under-utilizes the multi-faceted nature of real-world anomalies and the complementary semantic representations across different transformer layers. To address these limitations, we propose CRAG-CLIP, a source-only zero-shot anomaly detection framework that models abnormality as a cross-layer relational geometry in the CLIP text space. The framework represents abnormality through multiple object-agnostic normal–anomaly semantic relations, constrains the learned directions with frozen language priors to mitigate source-specific semantic drift, and promotes complementary semantic coverage through diversity regularization. It further estimates the reliability of layer-wise patch evidence from posterior uncertainty and cross-layer disagreement, enabling data-dependent fusion across visual layers. For image-level prediction, global semantic evidence and risk-sensitive local anomaly evidence are jointly aggregated to capture both holistic context and localized defects. Extensive evaluations on five industrial and nine medical datasets establish the effectiveness and competitive performance of CRAG-CLIP.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.