STAR: Practical Backdoor Attacks Against Vision-Language-Action Models via Semantic Triggers
Abstract
While Vision-Language-Action (VLA) models are increasingly deployed for real-world robotic control and open-ended tasks, their broader adoption also increases their exposure to physical backdoor attacks. Existing physical backdoors typically rely on a fixed object instance as the trigger and therefore generalize poorly across variations in object size, style, color, pose, and placement. Naively diversifying trigger instances does not resolve this limitation, as the model may instead learn an insertion-based shortcut that activates on arbitrary foreground objects. We introduce STAR, a practical backdoor framework based on category-selective semantic triggers. Diverse instances of a target object category activate the attack in trigger-present scenes, while benign behavior is preserved in clean and non-trigger scenes. STAR consists of two stages: Semantic Triplet Construction with Matched Object Interventions, which generates matched observations from clean, trigger-present, and non-trigger scenes under the same robot states and task contexts; and Category-Selective Trigger Injection via Triplet-Contrast Optimization, which associates malicious behavior with the target category while suppressing insertion-based shortcuts. Experiments on OpenVLA-OFT and demonstrate that STAR generalizes across distinct VLA architectures. Evaluations across the four LIBERO suites, downstream adaptation on MimicGen, and real-world manipulation with SO-ARM100 further demonstrate robust activation on diverse and previously unseen target instances while largely preserving benign task performance. These results reveal semantic object triggers as a broader and more practical backdoor threat to VLA models deployed in real-world environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.