acceptodds
Under review as a conference paper at ICLR 2027

From Rigid Carriers to Flexible Wrappers: Towards Output-Concealment Shadow-Activated Backdoor Attacks on MLLMs

Abstract

BadMLLM introduces the ***S***hadow-activated ***B***ackdoor ***A***ttack (***SBA***) for MLLMs. Malicious responses are triggered when generated content mentions a predefined target entity, rather than by explicit input triggers. (1) However, existing attacks, including SBA, neglect output-side concealment. Their rigid response patterns remain detectable by output-analysis defenses. This paper decomposes the backdoor carrier, i.e., the full sentence containing the malicious goal, into an invariant goal and flexible wrappers. The decomposition shifts similarity of backdoor carriers from sentence-level to lexical-level. (2) Unfortunately, directly applying Supervised ***F***ine-***T***uning (***SFT***)-based methods, such as BadMLLM, fails to lead victim MLLMs to associate entity mentions with goal insertion under varied wrappers. To resolve this tension, we formulate backdoor implantation as a constrained multi-objective optimization problem. Leveraging the flexible reward-driven optimization of ***R***einforcement ***L***earning (***RL***), a diverse-Wrapper SBA framework (***WrapperSBA***) is proposed to jointly optimize attack effectiveness, wrapper diversity, positional variation, and response quality. A sample-level instruction decay mechanism further addresses the association difficulty and gradually aligns training with prompt-free inference. (3) Conventional RL reward designs use fixed weights as hyperparameters to balance different reward components, which can cause reward hacking by overfitting to individual components. This paper introduces a dynamic weight scheduler that adaptively recalibrates coefficients according to measured contributions. Experiments show that WrapperSBA increases ASR from 46.57% to 99.40% and substantially reduces Self-BLEU of wrappers from 1.00 to 0.40. WrapperSBA substantially reduces the detection effectiveness of the evaluated defenses relative to fixed-wrapper SBA by enhancing output-side concealment. Further validation across unimodal settings and diverse attacks supports the defense-evasion benefit of wrapper diversity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.