What Success Rate Hides: Mapping Disturbance Tolerance in Vision-Language-Action Policies
Abstract
Vision-language-action (VLA) policies are typically ranked by task success, but success captures only the endpoint of behavior rather than how a policy responds during execution. Two policies that both complete a task from a nominal state may behave very differently when the same in-progress state is physically displaced. We measure this directly using a calibrated, policy-free push that produces a single mid-task release state, which is restored bit-identically for every policy. A disturbance-tolerance map then records next-milestone success across push directions and severities, while a scalar score, Qµ, summarizes performance under a declared disturbance distribution and the full map reveals where policy behaviors diverge. Across three released LIBERO policies, clean success does not resolve disturbance tolerance: 18 of 20 fresh anchors at which all policies are clean-competent separate under at least one matched disturbance, and all 28 selected decisive cells preserve their performance gap after expansion to 10 stochastic draws, including local reversals of the aggregate ranking. These differences persist even when execution cadence is equalized. VLA evaluation should therefore report clean competence, the tolerance score, and the disturbance-tolerance map, together with the execution setting under which they were obtained.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.