Doubt Before Acting: Assumption-Level Representation Editing for Safe Embodied LLM Planning
Abstract
Embodied agents increasingly use large language models (LLMs) to turn goals into plans with many steps, and an unsafe physical action often cannot be undone. A plan that executes correctly can still be unsafe when it relies on unchecked assumptions about the world, such as a corridor staying clear or a person staying where last seen. To act safely, an agent needs to know whether its plan still works when these assumptions break, and it needs to know this before the first action. We introduce DoubtBeforeAct, an inference-time layer that performs this check. The layer extracts typed assumptions from the scene and uses one catalog of typed operators in two modes. Propose edits the scene to obtain different candidate plans, and Doubt edits the assumptions to build plausible worlds in which they fail. DoubtBeforeAct executes every candidate in these worlds, picks the plan whose worst outcomes are best, and monitors its assumptions during execution. We evaluate on five embodied safety benchmarks and three long-horizon embodied benchmarks against thirteen baselines at the same budget of model calls and executions. DoubtBeforeAct raises safe and successful completion by 8.1 points and long-horizon success by 7.2 points with Qwen3-32B, and by 7.9 to 14.6 points with Qwen3-8B and a frontier model. Test worlds come from assumption classes withheld during development and from a separate authoring pipeline. At equal wall-clock time it still beats the strongest baseline on all eight benchmarks, and as a red teaming tool DoubtBeforeAct finds 205 avoidable failures per 1,000 runs against 104 for the best external generator.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.