AnchorCLIP: Enhancing CLIP Robustness through Interpretability
Abstract
Pretrained Vision-language models such as CLIP show impressive zero-shot performance on clean images. However, they are extremely vulnerable to common image corruptions, such as Gaussian noise, Pixelation or Motion blur. Existing approaches to improve robustness focus mostly on adversarial fine tuning but they have crucial limitations. They are ineffective for severe image corruptions, involve fine tuning the entire CLIP encoder, and are largely black box. In this paper, we leverage CLIP’s semantic interpretability to enforce robustness. We show that enforcing semantic consistency in concept space across corruptions significantly improves robustness. Our proposed method AnchorCLIP, uses a lightweight adapter to align the corrupted image embeddings to the interpretable latent space of clean CLIP embeddings, effectively “anchoring” the noisy representations to their semantic meanings. Extensive experiments indicate that AnchorCLIP outperforms existing adversarial robustness and interpretability based methods on widely used benchmark datasets, while maintaining performance on clean data. Distinct from previous work, AnchorCLIP is based on interpretability, is lightweight, and consistently outperforms strong baselines for severe image corruptions. By using interpretability as a mechanism for robustness instead of just post-hoc explanations, AnchorCLIP provides a scalable path toward vision-language systems that are both inherently robust and transparent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.