DA-VLA: A Dynamics-Aware Vision-Language-Action System for High-Speed Manipulation
Abstract
Existing VLA systems remain limited in task execution speed on real robots, making it difficult to meet the throughput demands of practical applications. To improve execution throughput while reliably completing tasks, a VLA system needs to naturally connect continuously updated policy outputs with ongoing motion, preserve motion stability and execution reliability, and fully exploit the robot's dynamics. To this end, we present DA-VLA, a dynamics-aware VLA system for high-speed manipulation that coordinates continuous policy action generation with the robot's actual motion capabilities. On the policy side, we introduce an execution-aware objective that promotes smooth motion within each action chunk during training while constraining continuity between newly generated chunks and already committed motion. The same objective guides action generation during inference, allowing successive policy updates to naturally extend the motion already being executed. On the execution side, DA-VLA evaluates the robot's available motion capability from its current state and continuously adapts execution according to upcoming action trends, avoiding unnecessary conservative slowdowns between successive policy targets and enabling fast, stable, and accurate motion within the robot's dynamic capabilities. Further speed-success comparisons show that, at a matched task success rate, \svla achieves nearly twice the execution throughput of . These results show that \svla can better exploit the robot's physical capabilities to improve practical execution throughput while preserving task reliability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.