Learning Beyond Bodies: Joint Learning and Transfer for Robust Humanoid Motion Tracking
Abstract
Conventional humanoid motion tracking approaches often condition control policies on retargeted robot motion, potentially carrying retargeting artifacts into motion execution. Beyond reference quality, a central question is how to share motion knowledge across robots and reuse it when adapting to a new embodiment with limited motion data. We address these challenges by separating shared human-motion conditioning from embodiment-specific control. Our motion encoder takes body-centric human keypoints as input, avoiding direct dependence on retargeted joint references while retaining retargeted trajectories for training supervision. On our primary motion benchmark, this input yields average tracking success and position accuracy comparable to robot joint-reference conditioning. Controlled single-embodiment comparisons additionally show smoother execution on this benchmark and selected cases with retargeting-induced joint discontinuities. This common input enables joint reinforcement learning of a shared encoder across robots, while dedicated control heads combine its features with proprioception and reference orientation. Across heterogeneous simulated humanoids, joint training improves position tracking on the same benchmark over corresponding single-robot training. This separation also supports adaptation: an embodiment excluded from pretraining reuses the encoder and learns a new control head. In a low-data setting, freezing the transferred encoder improves position tracking on an external motion dataset relative to training from scratch. These results demonstrate that our framework supports more robust motion execution, effective joint learning across humanoids, and the transfer of motion knowledge to new embodiments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.