acceptodds
Under review as a conference paper at ICLR 2027

Backdoor Success Outside the Data Manifold: A Geometric Study of CLIP Backdoor

Abstract

Contrastive Language–Image Pretraining (CLIP) has substantially advanced vision–language representation learning through self-supervised pretraining on web-scale image–text pairs. However, its heavy reliance on uncurated web data leaves it highly susceptible to backdoor attacks. Existing CLIP backdoor attacks can achieve high attack success rates, but they often leave a detectable geometric trace, where backdoor representations are pushed outside the clean data manifold, making them vulnerable to representation-based defenses. We mainly attribute this to a pervasive geometric phenomenon across CLIP backdoors, named backdoor representation off-manifoldness. In this work, we trace the phenomenon to the modality gap between CLIP's image and text embedding spaces because cross-modal backdoor objective pulls a triggered image(text) representation toward a textual(image) target, and a large fraction of the backdoor gradient lies outside the subspace supported by image(text) representations. Guided by this insight, we propose the On-Manifold Backdoor Attack (OMA), which retains the standard cross-modal backdoor objective while regularizing triggered representations against an intra-modal target feature bank via a principal-component normal loss and a local mean-squared-error loss. Across ImageNet-1K, ImageNet-R, and ImageNet-Sketch, OMA maintains a near- undefended attack success rate while effectively suppressing anomaly detection signals. Our study connects CLIP backdoor detectability directly to the modality gap and demonstrates that attack success can be decoupled from off-manifold detection signatures. Our code is in the supplementary material and will be open-sourced upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.