Outcome or Process? Rethinking Supervision for Flow-Matching VLA Policies
Abstract
Flow-matching vision-language-action (VLA) policies generate actions through a solver trajectory, repeatedly evaluating a conditional velocity field at its intermediate states. This raises a fundamental question for policy distillation: should a student imitate a sampled teacher action, or the field that generates it? Distilling the teacher’s final action chunk yields conditional flow-matching targets, but a single action chunk does not uniquely specify the teacher’s local field or its intermediate computation. We study this choice with Path-wise Velocity Distillation (PVD), which regresses the frozen teacher’s velocity at detached states along the student’s own solver trajectory. Under matched teacher–student checkpoints and training exposure, PVD outperforms DAgger with action-chunk labels by 9.9 and 10.1 percentage points on ManiSkill at four- and eight-step deployment, 12.0 points on Meta-World, and 0.180 subtasks per chain on CALVIN. A one-seed ManiSkill ablation attributes more of the gain to direct teacher-velocity targets than to placing them on the student’s solver trajectory. The method fails when trained and deployed at two solver steps. Under this matched setup, velocity labels on the student’s own solver states yield better students than action-chunk labels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.