PUARobot: Psychological Manipulation for Jailbreaking Embodied LLM Agents
Abstract
Embodied AI agents increasingly rely on large language models (LLMs) for language understanding and task planning in real-world tasks, yet LLM jailbreak vulnerabilities may translate harmful text into dangerous physical actions. Inspired by psychological manipulation in human interactions, we ask whether similar social, emotional, and cognitive pressure can bypass embodied-agent safeguards and induce unsafe actions. In this work, we investigate psychological manipulation-based jailbreak attacks on embodied agents. Specifically, we (1) propose PUARobot, a black-box jailbreak attack framework that systematizes seven psychological-manipulation strategies as attack methods to reveal the safety weaknesses of embodied agents under different forms of manipulative pressure; (2) establish a paradigm that transforms a direct single-turn request into an attacker-authored, multi-turn, psychological-manipulation-based disguised dialogue, gradually applying pressure in the context to induce unsafe action plans; and (3) develop an action-centered embodied-safety evaluation mechanism and feed evaluation feedback into dialogue search to iteratively optimize attack prompts, addressing the tendency of existing benchmarks to overemphasize text and misjudge unsafe actions. Experiments across multiple open- and closed-source models show that PUARobot achieves an average Attack Success Rate (ASR) of 76.2%, compared with 7.3% for the strongest BadRobot baseline. Further experiments on Code as Policies, ProgPrompt, and VisProg demonstrate its effectiveness across mainstream embodied AI frameworks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.