Beyond Feasibility: Enforcing Safety Rules on Language-Model Planners with World-Aware Look-Ahead
Abstract
Language models can turn natural-language instructions into robot plans, but their plans carry no safety guarantee: a plan can be executable and reach its goal yet still violate safety rules. When the rules and the action model are formally specified, one would like a planner that never returns an unsafe plan, preserves the language model’s preferences among safe plans, and needs only a single decoding pass. We introduce GUARDCTRL, which weights each candidate token by the probability that the model completes it into a safe plan. This look-ahead tracks the world state and the progress of each rule along every hypothetical continuation, and is computed exactly with a hidden Markov model distilled over grounded action and object symbols. Its support guarantees that one decoding pass returns a safe plan whenever one exists within the token horizon, and its magnitude preserves the model’s preferences among safe plans. We evaluate on the DESPITE benchmark and on a new, harder set with rules over the action history. With Llama-3.1-8B-Instruct, GUARDCTRL is safe on 98.4% of DESPITE hard tasks and 98.2% of tasks with temporal rules, against 13.8% and 6.4% for rejection sampling with up to ten attempts, and close to GPT-5 given the same rules. Its extra time, mostly a one-time compilation, is about twice that of one unconstrained generation. It also follows preferences stated only in the instruction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.