acceptodds
Under review as a conference paper at ICLR 2027

RoboSafeBench: Diagnosing Dynamic Contact Risk in Vision-Language-Action Policies

Abstract

Recent progress in Vision-Language-Action models has demonstrated strong generalization across diverse manipulation tasks, yet their safety behavior is still primarily evaluated in static, human-free environments. This poses a critical deployment gap: real environments are dynamic, and robots must react to moving obstacles and human intrusions while preserving both success rates and physical safety. We introduce RoboSafeBench, a non-invasive diagnostic meta-benchmark that upgrades existing robot manipulation benchmarks into dynamic contact testbeds while preserving each host task’s instruction, observation–action interface, and success criterion, and adding only a scheduled intruder during evaluation. RoboSafeBench is a measurement tool, not a safety certificate: it records task success and robot–intruder contact jointly, separating contact-free from contact-involved success. We define four evaluation levels—static execution, sudden object intrusion, sliding object intrusion, and human-like reaching intrusion—and instantiate them on LIBERO, RoboTwin, and SimplerEnv across 39 task configurations. We find that: (i) high static success does not imply low dynamic contact; (ii) under human-like reaching, contact-involved success is the most common outcome for every evaluated policy, even where task success stays within 0.7 points of the static rate; (iii) a success-rate lead can consist almost entirely of contact-involved success: under sudden intrusion on SimplerEnv, XVLA exceeds StarVLA by 17.8 points of success rate but by only 0.4 points of contact-free success rate. A privileged yield-to-home reference for on LIBERO lowers reaching-level contact from 60.8% to 3.7% while retaining 89.8% success, showing that such an operating point exists on this host and is not reached by the evaluated policies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.