Bridging Vision-Language-Action Models and Whole-Body Controller via Structured Action Interfaces
Abstract
Quadrupedal loco-manipulation systems remain underexplored. Extending humanoid vision-language-action (VLA) pipelines is not straightforward: this composite embodiment has no direct kinematic counterpart from which motion references can be retargeted, raising two open questions: how an arm-body-coordinated whole-body controller should be derived, and how it should be bridged with a VLA policy using an appropriate action representation. In this paper, firstly, we propose an arm-body-coordinated whole-body controller to serve as an embodiment-specific motion generator. The use of such a native tracking interface inevitably introduces irreducible tracking errors, making it unsuitable for direct VLA integration. Thus, we transfer this generated coordination into a structured controller with substantially lower error, coordinating the upper and lower body explicitly through the degrees of freedom at the arm mount, and we subsequently learn a VLA policy using this structured interface. The resulting system, QuadArm-VLA, spans a wide range of mobile, contact-rich, and constrained manipulation tasks on a Unitree Go2 quadruped equipped with an ARX X5 arm. Our results provide practical guidance for acquiring coordinated motion and constructing VLA action interfaces for embodiments without a natural kinematic counterpart, and we will open-source our system codes to aid future research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.