GeoPhase: Phase-Aware Geometric Safety Filtering for Vision-Language-Action Policies
Abstract
Vision-language-action (VLA) policies enable general-purpose robot manipulation, but their task-directed actions lack explicit geometric safety guarantees, leaving the robot vulnerable to physical collisions in cluttered environments. While runtime safety filtering via control barrier function quadratic programs (CBF-QPs) offers an interpretable solution, current approaches are bottlenecked by costly multi-stage geometric perception, phase-agnostic constraint scheduling, and insufficient protection scopes that typically ignore the grasped object and arm body. To address these limitations, we present GeoPhase, a runtime geometric safety layer that operates on a hybrid learned-perception and analytic-control design. GeoPhase bypasses explicit perception pipelines by predicting structured obstacle ellipsoids directly from intermediate VLA visual embeddings. By tracking the manipulation phase (approach, grasp, transport, place) online via proprioceptive and gripper signals, it dynamically adapts constraint activation to optimize the phase-specific balance between safety and task completion. Furthermore, GeoPhase constructs a multi-proxy geometric representation of the robot-object system, encompassing the end-effector, held object, and arm body, to jointly constrain the robot–object geometry via a unified, minimal-deviation CBF-QP action projection. On the SafeLIBERO benchmark, GeoPhase achieves an average collision rate of and a safe success rate of , outperforming the best reported baseline averages by and percentage points, respectively. These results demonstrate that GeoPhase substantially improves collision-free task completion without modifying the underlying, frozen VLA policy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.