CrossCLIP: Bidirectional Cross-Modal Interaction with Instance-Adaptive Frequency Modulation for Anomaly Detection
Abstract
Universal anomaly detection aims to identify anomalies in unseen object categories and domains without target-specific training data. Recent methods leverage pretrained vision-language models, such as CLIP, to align visual and textual representations of normality and abnormality. However, they typically adapt the two modalities independently or rely on unidirectional interaction, limiting the use of complementary cross-modal information. To address this limitation, we propose CrossCLIP, which enables bidirectional cross-modal interaction through three adapters: a Vision-Conditioned Textual Adapter (VCTA), a Text-Conditioned Spectral Adapter (TCSA), and a Textual Reference Adapter (TRA). VCTA incorporates image information into text prompts to generate instance-specific normality and abnormality anchors, which in turn guide TCSA to perform instance-adaptive frequency modulation for capturing transferable anomalous patterns across diverse objects and domains. TRA further generates a text-grounded normality reference to complement normal-image-based comparison in zero- and few-shot settings. Extensive experiments on diverse industrial and medical datasets demonstrate that CrossCLIP consistently outperforms the state-of-the-art universal anomaly detection methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.