acceptodds
Under review as a conference paper at ICLR 2027

When Text Dominates: Understanding and Designing Dual-Key Backdoors in MLLMs

Abstract

Dual-key backdoors in multimodal large language models (MLLMs) aim to elicit an attacker-chosen response only when image and text triggers appear together. We investigate whether this joint requirement can be learned with low-amplitude universal adversarial perturbations (UAPs) that limit visible image changes. We find that activation can become dominated by the text trigger, even when single-trigger inputs are trained to receive normal responses. Through gradient analysis and visual-state replacement, we find that the UAP in this text-dominated setting produces substantial changes in hidden states, yet these changes have little influence on the target response. We call this a representation-influence mismatch. A path-integral decomposition explains why large representation changes alone do not ensure influence on the target response. Guided by this analysis, we construct UAPs from normalized projector residuals and, under training control, add a necessity penalty to strengthen dependence on both triggers. Across datasets, model architectures, and fine-tuning settings, our method achieves high attack success when both triggers are present and low activation with either trigger alone, while maintaining clean-task performance comparable to clean fine-tuning. Ablations support the trigger construction, while visual-state interventions show that the learned visual trigger contributes to the target response.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.