ProtoLoRA: Visual-Prototype-Guided LoRA Backdoor Attack against Vision–Language Models
Abstract
Parameter-efficient adaptation turns a pretrained vision–language system into a composition of a frozen backbone and independently obtained parameter updates, creating a component-level supply-chain risk that is not captured by conventional full-model backdoors. We study whether a compact third-party LoRA adapter can impose selective semantic control over a frozen contrastive representation space while preserving expected behavior on clean inputs and without access to the victim's downstream data, prompts, candidate labels, or task configuration. We formulate this setting as trigger-conditioned representation steering under a low-rank parameter budget. ProtoLoRA combines target-related visual prototypes for capacity-aware semantic relocation, cross-modal target-priority ranking against competitive texts, and a low-distortion Lab-DCT trigger for conditional activation, together with clean contrastive regularization. Across zero-shot classification and image–text retrieval, the resulting adapters achieve high targeted redirection with limited clean-utility degradation and small input-space distortion, showing that lightweight adaptation modules can constitute a security-critical control surface even when the pretrained backbone remains frozen.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.