Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots
Abstract
We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel grippers. Benefiting from the robot-free setting, human action data becomes one of the most scalable sources for manipulation data. However, transferring manipulation skills from humans to robots remains hard: most prior work treats humans as just another bi-manual 6DoF embodiment, where hand-pose estimates are noisy and the contact patterns of human fingers differ fundamentally from those of a parallel gripper. We argue that learning rotation-inclusive action signals from human data is therefore sub-optimal and instead propose a bridging action representation, *i.e.*, the relative wrist translation within the head-camera frame, an action space shared by humans and robots. To handle the potential absence of the 6DoF end-effector action or gripper signal in human data, we build a -like vision-language-action model with interleaved action tokens. On a suite of bi-manual manipulation tasks seen only in human action data, our bridging action enables a superior success rate of 22.50% over naively adopting noisy 6DoF actions 12.50%, and can be further improved to 38.33% with large-scale human-only pre-training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.