WaypointVLA: Training-Efficient Vision-Language-Action Learning via Sparse Action Waypoints
Abstract
Vision-language-action (VLA) policies typically predict dense action chunks over a long, fixed future horizon, requiring supervision for every intermediate action. However, under closed-loop control, only a small portion of each predicted chunk is executed before replanning, creating a prediction–execution mismatch. Our analysis of a policy on LIBERO further shows that distant future actions become increasingly unreliable, questioning the necessity of dense long-horizon supervision. Meanwhile, within the practically useful horizon, action trajectories can be effectively reconstructed from only a sparse set of key actions using simple linear interpolation. Motivated by these, we introduce WaypointVLA, which predicts only two key actions—a start-point action and an endpoint action—while lightweight interpolation reconstructs the intermediate low-level controls. To support training, we derive sparse waypoint supervision directly from existing demonstrations and use it to fine-tune the original VLA model without modifying the base VLA backbone. By concentrating supervision on temporally informative actions rather than redundant intermediate controls, WaypointVLA achieves faster convergence and improved robustness to visual and environmental perturbations. We evaluate WaypointVLA on LIBERO and LIBERO-Plus using as the base model. For clarity, we refer to the original baseline as the dense policy. Remarkably, using only 200 demonstrations (13.9% of the full-data setting), WaypointVLA surpasses the full-data dense policy on both benchmarks while achieving a training speedup. These results suggest that sparse waypoint training provides a more efficient alternative to dense training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.