Learning Expressive and Compositional Motion Representation via Spectral Skills
Abstract
Robotic foundation models offer a promising path toward general-purpose humanoid robot control often with hierarchical architectures. However, their effectiveness depends on a command interface between the planner and the controller, which must support accurate execution while remaining easy to predict, and ideally compose new behaviors from prior ones. In this work, we introduce a latent-representation of such interface we term spectral skills, which meets these requirements through predictive representation learning. By design, the spectral skills compactly encodes short motion segments and are learned by predicting subsequent motion rather than reconstructing the encoder’s input. On a 29-DoF humanoid, a controller conditioned on spectral skills reduces global tracking error by relative to the state-of-the-art. The same frozen controller chains independently encoded skills without a separate transition policy. It also composes new behaviors by adding orthogonal directions to any compatible base skill, producing combinations that are unseen from the training data. We demonstrate tracking, chaining and composition, as well as through a languaged conditioned planner on Unitree G1 hardware. Project page: https://spectral-skill.github.io.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.