acceptodds
Under review as a conference paper at ICLR 2027

Beyond Feasibility: Enforcing Safety Rules on Language-Model Planners with World-Aware Look-Ahead

Abstract

Language models can turn natural-language instructions into robot plans, but their plans carry no safety guarantee: a plan can be executable and reach its goal yet still violate safety rules. When the rules and the action model are formally specified, one would like a planner that never returns an unsafe plan, preserves the language model’s preferences among safe plans, and needs only a single decoding pass. We introduce GUARDCTRL, which weights each candidate token by the probability that the model completes it into a safe plan. This look-ahead tracks the world state and the progress of each rule along every hypothetical continuation, and is computed exactly with a hidden Markov model distilled over grounded action and object symbols. Its support guarantees that one decoding pass returns a safe plan whenever one exists within the token horizon, and its magnitude preserves the model’s preferences among safe plans. We evaluate on the DESPITE benchmark and on a new, harder set with rules over the action history. With Llama-3.1-8B-Instruct, GUARDCTRL is safe on 98.4% of DESPITE hard tasks and 98.2% of tasks with temporal rules, against 13.8% and 6.4% for rejection sampling with up to ten attempts, and close to GPT-5 given the same rules. Its extra time, mostly a one-time compilation, is about twice that of one unconstrained generation. It also follows preferences stated only in the instruction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.