Uncovering Vulnerability of Vision-Language-Action Models under Joint-Level Physical Faults
Abstract
Robustness studies of Vision-Language-Action (VLA) models have largely focused on perceptual, task, and command-space perturbations, while changes in how the robot body physically realizes actions remain comparatively underexplored. We investigate this embodiment-side robustness gap through joint-level physical faults, including joint locks, range restrictions, and increased friction. These faults alter the physical conditions of action execution without directly perturbing policy outputs, introducing command–execution mismatch into the closed loop. Our evaluation of OpenVLA-OFT and \(\pi_0.5\) on LIBERO reveals substantial degradation that varies strongly with the affected joint and fault condition. Joint-dependent vulnerability is related to nominal joint usage and end-effector sensitivity, while increased friction reduces task success even without narrowing joint-position limits. We examine complementary mitigation strategies with the VLA backbones frozen: language-based fault feedback and fault-aware local kinematic correction provide no consistent task-level gains in the evaluated settings, whereas embodiment-side calibration partially mitigates degradation across both backbones while largely preserving fault-free performance. Together, our findings highlight changes in physical execution as an important dimension of VLA robustness that requires explicit evaluation alongside perceptual and task variations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.