acceptodds
Under review as a conference paper at ICLR 2027

LIBERO-Risk: Safety-Critical Scenario Generation and Process-Level Evaluation for Vision-Language-Action Models

Abstract

Vision-Language-Action models (VLAs) map visual observations and language instructions to robot actions. Successful task completion, however, does not necessarily imply safe execution. Evaluating VLA safety therefore requires introducing meaningful risks into source tasks while preserving their semantics and executability. Automating this process is challenging since naive risk injection can produce invalid scenes, alter source-task semantics, or render the task unexecutable. In this paper, we present LIBERO-Risk, the first framework for automated safety-oriented scenario construction in VLA evaluation. Through staged verification and iterative revision, LIBERO-Risk preserves source-task semantics and executability while refining invalid candidates into valid risk scenarios. We further introduce a multidimensional process-level scoring strategy that evaluates VLA safety through contact, spatial, and temporal signals. Experiments show that LIBERO-Risk succeeds in 174 of 234 (74.4%) construction trials and in 64 of 78 (82.1%) task–mutation families. Evaluations across six VLAs further expose safety-relevant behavior that task success alone does not capture.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.