acceptodds
Under review as a conference paper at ICLR 2027

ManiGuard: A Specification-Grounded Framework for Safety Evaluation and Demonstration Generation in Robotic Manipulation

Abstract

Foundation-model policies for robotic manipulation are rapidly improving in task success, yet whether they complete tasks safely remains poorly characterized. We introduce ManiGuard, a framework that provides safety evaluation for manipulation policies and safety-annotated demonstrations for downstream policy learning, both grounded in the same formal specifications, expressed in linear temporal logic over finite traces (LTLf). Its benchmark, ManiGuard, contains 200 contact-rich household tasks across six families, with safety defined independently of task success; one in-distribution condition and four single-axis out-of-distribution perturbations per task yield 1,000 fixed scenarios in physics simulation, each checked at every step of a rollout by automaton monitors compiled from the LTLf specifications over physics-grounded predicates rather than learned judges. Its demonstration pipeline uses the same monitors to verify and annotate trajectories, yielding 8,000 safe, successful demonstrations (40 per task) for supervised fine-tuning (SFT). Across more than 23,000 rollouts of vision-language-action policies, task success remains an unreliable proxy for safety: 6–21% of successful SFT rollouts incur a counted safety violation. Standard SFT on the generated demonstrations raises safe-success rates to 7.5–29.8% and improves engagement-conditioned safety relative to available zero-shot counterparts, yet 21–42% of engaged SFT rollouts still violate the constraints, and failures persist under distribution shift and on a physical Franka running matched tasks. ManiGuard thus links specification-grounded data collection, standard policy learning, and safety measurement in one workflow, and exposes the safety gaps that remain.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.