acceptodds
Under review as a conference paper at ICLR 2027

Hidden in Plain Pieces: Jailbreaking LVLM-based Embodied Agents via Compositional Environment Injection

Abstract

Large Vision-Language Models (LVLMs) are widely integrated into embodied agents to improve multimodal understanding and planning capabilities, and recent research has increasingly explored jailbreak attacks to reveal their vulnerabilities. However, existing methods have not systematically investigated how multimodal environments affect the safety of embodied LVLMs, particularly when seemingly benign environmental cues combine to convey malicious intent. To bridge this gap, we propose a black-box, generalizable, and easily deployable jailbreak method termed CoEnv, the first compositional environment injection attack that writes decomposed malicious queries in the environment to compromise embodied LVLMs. Specifically, three vulnerabilities are exploited to achieve our attack across perception, cognition, and action stages: (i) multimodal perception confusion, (ii) structure-biased cognition, and (iii) imbalanced action safety alignment. Furthermore, we construct SafeEnvBench, the first benchmark to evaluate embodied LVLMs’ safety across diverse multimodal environments, with human-curated, high-quality samples. Based on our benchmark, extensive experiments, including both simulator and physical world scenarios, demonstrate the effectiveness and superiority of our proposed attack against state-of-the-art competitors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.