Flow Matching Is Not a Contractive Class for Vision-Language-Action Policies
Abstract
Vision-language-action (VLA) policies are increasingly employing flow matching for robot action generation. The process starts from random noise and produces an action over a few steps. It is currently believed that flow matching intrinsically contracts noise perturbations and this leads to robust VLA policies. Recent approaches often build upon this idea. We show that this belief is actually not correct. Contraction and expansion of flow in VLAs are mainly tied to the model architecture, not to the flow matching process. We establish this with a rigorous treatment of four diverse state-of-the-art VLAs. We achieve a clear split of two contractions and two expansions, which is not caused by flow matching, the direction of the flow, or the number of steps. Two policies sharing similar process design choices still land on the opposite sides, isolating the action-generation module as the primary cause. Our analysis further reveals two crucial insights. (a) Contraction in a deployed policy must be judged over the entire flow, not one step at a time, which is the current norm and can be misleading. We provide results supporting expansion at every step for a policy whose flow as a whole still contracts. (b) For flow-matching VLAs, robustness must not be inferred from the flow and needs to be measured independently because model robustness may not align with contraction. Our supporting evidence identifies a flow expanding policy as the most robust one. Our source code will be made public after acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.