PlanSmuggle: Candidate-to-Trajectory Risk in Embodied LLM Planning
Abstract
Embodied LLM agents increasingly rely on language models to synthesize executable plans from natural-language goals and multiple candidate branches. Prior embodied-agent jailbreaks mainly target explicitly unsafe prompts or individual primitive actions, leaving plan synthesis comparatively understudied. We introduce , a black-box attack on the plan-synthesis stage: it exposes a certified risky candidate during branch synthesis and induces the final executable plan to reproduce its violating transition structure without requiring an explicitly unsafe user goal. In a successful attack, every atomic action passes schema and immediate-precondition checks, yet the trajectory produced by executing their sequence violates a global safety invariant. We formalize this compositional failure and introduce PlanSmuggleBench, comprising 120 symbolic planning tasks across four constraint families. Across Qwen3-32B, DeepSeek-V4-Pro, GPT-5, and Qwen3.7-Plus, PS-Full achieves verified ASRs of 100.0%, 80.0%, 94.2%, and 75.8%, respectively. In a separate paired risk-by-merge intervention, attacks occur only when risky-candidate exposure and merge permission are jointly enabled (45.8%; 0% in each remaining cell), providing controlled evidence for their interaction. Under the benchmark's deterministic transition model, exact trajectory verification blocks every observed attack while preserving benign task completion. These results establish candidate-to-trajectory synthesis as a distinct safety boundary for embodied LLM agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.