Robustness Margins for Local Visual Perturbations in Vision–Language–Action Policies with Flow Matching
Abstract
Vision-language-action (VLA) policies often remain effective under visual disturbances such as changes in lighting, but it is not well understood why their behavior is insensitive to some perturbations and highly sensitive to others. As a first step toward understanding this robustness, we study a more local question: given a fixed observation and a perturbation to its visual conditioning, how much can the distribution of executed action chunks change? We formulate this question probabilistically and derive a local bound on the shift of the action distribution under small perturbations. The bound depends on the magnitude and direction of the input perturbation together with the sensitivity of the policy’s flow to its conditioning variable, characterized through the Jacobian. This provides a way to relate perturbations in observation space to changes in the distribution of sampled actions without making assumptions about downstream task success. We empirically study this relationship in LIBERO using several classes of visual perturbations and compare the predicted sensitivity with the observed shifts in action distributions. Our results provide an initial framework for analyzing VLA robustness through the geometry of how conditioning perturbations propagate through the generative action policy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.