PRESTO: Prioritized Reasoning for Safe Trajectory Optimization with Foundation Models
Abstract
Ensuring safety in open-world embodied systems is difficult not only because novel objects and hazards arise at test time, but because existing formal methods treat all safety constraints uniformly–leading to overly conservative or infeasible behavior in complex scenes. This paper introduces PRESTO, a framework that leverages multimodal large language models (MLLMs) to generate interpretable safety constraints directly from visual observations, and embeds that into a prioritized safe planner. Our framework enables (i) open-vocabulary safety reasoning over previously unseen objects, (ii) priority-aware safety constraints based on semantic risk, (iii) planner-agnostic deployment as a safety layer, and (iv) interpretable safety constraints. We evaluate PRESTO across autonomous driving and robotic manipulation tasks, reducing collision rates by 59.25% when applied as a safety layer to existing planners, while also generalizing to manipulation settings requiring semantic reasoning about fragile, hazardous, and low-risk objects. Our results highlight the potential of combining high-level multimodal reasoning with formal safety guarantees for scalable and interpretable autonomy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.