acceptodds
Under review as a conference paper at ICLR 2027

When the Safety Layer Is Not Safe: In-depth Analysis of VLA Failure Detection under Distribution Shift

Abstract

Vision-language-action (VLA) policies achieve impressive success rates on manipulation benchmarks, yet recent robustness analyses show this competence is brittle and collapses under modest open-world variation. Runtime failure detectors are increasingly used alongside VLA policies. Like the policies they monitor, these detectors perform well under ideal conditions, raising an alarm in time for the robot to stop, back off, or ask for help. Existing work, however, trains, calibrates, and evaluates these detectors only on clean rollouts and rarely tests them under the distribution shifts (e.g., a changed camera viewpoint or robot initial state) that cause policies to fail in the first place. We therefore collect a 5,501-rollout corpus of successes and failures of two VLA policies under seven controlled perturbation dimensions and introduce a three-distribution protocol that separates a detector's ranking ability from the validity of its calibrated threshold. We find that thresholds break before rankings: clean-calibrated detectors retain most of their discrimination under shift but read unfamiliar success as failure, raising false alarms by 16.5 percentage points on the shifts to which policies are most sensitive. The direction of the break is set by how the score accumulates signed per-step evidence, not by how far features drift: on the same features, kNN detectors over-alarm while a Mahalanobis detector falls silent and misses failures. Task novelty, by contrast, mainly erodes ranking. No evaluated repair restores nominal calibration without sacrificing detection, and label-free recalibration is circular because the unlabeled deployment pool contains the failures the monitor is meant to catch. Neither a high AUROC nor a low false-alarm rate therefore certifies a monitor. A safety layer built for the moment a policy cannot be trusted must be evaluated and calibrated under the shifts that create that moment, with scores that do not mistake novelty for failure. Code and data will be released upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.