GENLOCK: Towards Multimodal Conjunctive Backdoors in Large Vision–Language Models
Abstract
Large vision–language models (LVLMs) have seen rapid adoption across diverse domains, yet reliance on third-party adaptations introduces severe backdoor risks. In context-dependent scenarios, an adversary requires backdoor activation to occur strictly when a designated visual cue and question co-occur, preventing premature exposure under isolated conditions while retaining benign task answers. However, existing multimodal backdoors readily degenerate into single-modality shortcuts and disrupt task utility once activated. We present GenLock to realize covert targeted backdoors with interpretable multimodal conjunctive control. Our analysis shows that high joint activation often masks single-key shortcuts, and short payload prefixes can trigger payload continuation even when a key is absent. GenLock enforces joint dependency throughout generation via a ReLU-gated visual adapter alongside a Dual-Key Interaction Contrastive Loss, while Payload–Utility Decoupled Training safeguards host task utility through span-balanced supervision. Across five LVLM backends, GenLock is the only approach maintaining over 90% selectivity at near-clean utility. On LLaVA-1.5, joint activation reaches 100.00% with a worst-single-key leakage of only 0.33%. Conditioned on the initial payload token, full-sequence contrastive control slashes worst-single-key leakage from 83.92% to 0.42%, supporting conditional control throughout generation. GenLock also exhibits the highest attack survival (80%+) across most evaluated input-side defenses, establishing a controlled red-team benchmark for multimodal safety.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.