Distinguishing Repairable Deployment Failures from Vision-Language-Action Failures Before Rollout
Abstract
Vision-language-action policies that work in the lab often fail when the camera is bumped or the arm starts somewhere new. We show that most of these failures are not the policy's fault, and that a robot can tell which ones it can fix before running the policy at all. The robot uses its own body as a measuring instrument. Its encoders say where its arm should appear in the image, and comparing that with where it does appear reveals how the camera moved. From this the robot re-renders the view the policy was trained on, predicts how much of that view it can restore, and checks whether its own measurement can be trusted. On LIBERO-Plus with three policy families, the correction repairs 61 of 100 camera failures while breaking 18 of 140 successes (p = 1.3e-6), resetting to the training start repairs all 47 starting-pose failures on untuned suites, and on a held-out policy a pre-registered test of the correction with a refined pose passes (15 repaired, 3 broken, p = 0.004). Before any rollout, the fraction of the training view the correction restores ranks which failures it will fix, even within a single perturbation category where benchmark labels carry no information, and the calibration's own residual flags every failed measurement on the held-out policies. Tracing each remaining failure to its cause, at most 8 of 100 camera failures of the development policy belong to the policy, and the rest are failed measurements or views the camera can no longer see. We prove a minimax floor on calibration from the body alone. All experiments are in simulation. With learned masks, calibration reaches about a centimetre, but the closed-loop gain over the raw policy is not yet significant, and on real images mask quality limits calibration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.