WAVE-VLA: A Whole-Body Humanoid VLA with Executable Vision–Motion Codes Learned from Heterogeneous Human and Robot Data
Abstract
Human video–motion offers behaviors beyond scarce humanoid demonstrations, but embodiment and supervision differ across sources. We present WAVE-VLA (Whole-body Action Vocabulary for Execution), which learns from task-unpaired G1 demonstrations and egocentric human video–motion. Both sources enter a common G1 action space and shared vision–motion codebook to extend one policy's repertoire while retaining robot skills. Current RGB, state, and language select body and hand codes that decode into synchronized references. A scene-conditioned Action Expert predicts a residual from the body reference, which zero residual preserves exactly. On HE and Nymeria development cohorts, refinement reduces mean forward-kinematics error relative to decoded references by 21.1% and 15.2%, respectively. The same frozen body–hand policy completes 39 of 60 real-G1 trials across four HE-covered and two Nymeria-covered tasks without task-specific parameter updates or retrieved trajectories. The HE-only baseline scores 27/40 on HE-covered and 0/20 on Nymeria-covered tasks, versus 25/40 and 14/20 for the joint policy in the reported robot trials.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.