From Rigid Carriers to Flexible Wrappers: Towards Output-Concealment Shadow-Activated Backdoor Attacks on MLLMs
Abstract
BadMLLM introduces the ***S***hadow-activated ***B***ackdoor ***A***ttack (***SBA***) for MLLMs. Malicious responses are triggered when generated content mentions a predefined target entity, rather than by explicit input triggers. (1) However, existing attacks, including SBA, neglect output-side concealment. Their rigid response patterns remain detectable by output-analysis defenses. This paper decomposes the backdoor carrier, i.e., the full sentence containing the malicious goal, into an invariant goal and flexible wrappers. The decomposition shifts similarity of backdoor carriers from sentence-level to lexical-level. (2) Unfortunately, directly applying Supervised ***F***ine-***T***uning (***SFT***)-based methods, such as BadMLLM, fails to lead victim MLLMs to associate entity mentions with goal insertion under varied wrappers. To resolve this tension, we formulate backdoor implantation as a constrained multi-objective optimization problem. Leveraging the flexible reward-driven optimization of ***R***einforcement ***L***earning (***RL***), a diverse-Wrapper SBA framework (***WrapperSBA***) is proposed to jointly optimize attack effectiveness, wrapper diversity, positional variation, and response quality. A sample-level instruction decay mechanism further addresses the association difficulty and gradually aligns training with prompt-free inference. (3) Conventional RL reward designs use fixed weights as hyperparameters to balance different reward components, which can cause reward hacking by overfitting to individual components. This paper introduces a dynamic weight scheduler that adaptively recalibrates coefficients according to measured contributions. Experiments show that WrapperSBA increases ASR from 46.57% to 99.40% and substantially reduces Self-BLEU of wrappers from 1.00 to 0.40. WrapperSBA substantially reduces the detection effectiveness of the evaluated defenses relative to fixed-wrapper SBA by enhancing output-side concealment. Further validation across unimodal settings and diverse attacks supports the defense-evasion benefit of wrapper diversity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.