CLASS: Clustered Latent Actions Separated from Body and Scene for Action Generation
Abstract
Action representations shared across bodies would let skills learned on one robot be reused on another. Existing latent-action learning, however, entangles motion meaning, what is done, with body-specific dynamics, on which body it is done, in one latent variable, so cross-body reuse of the representation remains limited. We propose Clustered Latent Actions Separated from Body and Scene (CLASS), a framework for action generation that draws latent actions into one cluster per motion category across bodies through supervised contrastive learning and conditions the action decoder on body information as a domain embedding , or on the object position as a scene variable in manipulation. The center of each cluster is then a drivable signal shared across bodies and scenes. On four DMC and Loco-MuJoCo bodies, per-step latent actions and centroids drive locomotion within and across bodies and benchmarks. Other bodies' centroids drive held-out body-motion combinations, and adapting about 32 parameters adds a new body. Cross-body cumulative reward drops about 80% without , and about 88% with contrastive learning removed. With a continuity regularizer, cosine interpolation between centroids lowers the fall rate at mid-rollout switches from about 98% to 13 to 19%. In object manipulation in simulation, fixed per-action centroids drive reach, pick, place, and push across object positions, and their four-action mean success rate of 85.5% exceeds the 79.0% of per-step re-estimation. The separation thus lets a centroid computed on one body drive the same action on another body or scene, and an action can be specified by a category name, one unlabeled demonstration, or interpolation between centroids.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.