Dual-Expert Flow Matching for Dynamic Robot Manipulation
Abstract
Vision-Language-Action (VLA) models predict where a robot should move, but not how fast. The velocity profile is neither supervised nor predicted; the low-level controller reconstructs it from position targets. For quasi-static manipulation this is harmless. For dynamic manipulation, in which the robot tosses, flips, slides, or pushes, the velocity at contact decides the outcome, and position tracking distorts it. We show that accurate position does not imply accurate velocity and we establish the need to learn velocity from the demonstrations. We propose VIPER, which adds a velocity expert to a pretrained VLA and denoises velocity jointly with the VLA’s position expert from shared noise to ensure temporal alignment between the noise-free position-velocity actions. VIPER leaves the VLA’s backbone frozen and attaches to any VLA with a flow-matching action expert. We evaluate VIPER on three VLA backbones and six dynamic manipulation tasks in simulation and on a real robot, from tossing and flipping to catching and handover. VIPER raises mean success from 32% to 51%, and improves backbones that already reason about temporal context. Our project website is available at https://viper-robot.github.io/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.