Generative Action Synthesis: Inferring Robot Actions from Kinematic Demonstrations
Abstract
For dexterous hands, kinematic demonstrations from motion-capture gloves and DexUMI-style exoskeletons are easier and faster to collect than demonstrations collected via robot teleoperation, but they record only observed hand motion, not the corresponding joint commands. This missing information is crucial during contact, when applying forces to objects or the environment requires joint commands to differ from the hand's observed configuration. Robot teleoperation data directly capture this missing information through the residual between commanded joint targets and observed hand states. In this work, we propose Generative Action Synthesis (GAS), which trains a conditional generative model of these residuals from teleoperation data and uses it to synthesize action labels for kinematic-only demonstrations. For each frame, the model uses past and future visual observations together with observed joint states to sample plausible command residuals for relabeling kinematic-only demonstrations. GAS synthesizes useful actions zero-shot across diverse manipulation tasks beyond its teleoperation training distribution, including mug grasping, towel folding, cap twisting, test tube insertion, USB unplugging, and Rubik's cube manipulation. With these synthesized actions, we train policies for mug grasping, test tube insertion, cap twisting, and towel folding that achieve high success rates and outperform both naive state-as-action baselines and constant heuristic offsets. These results suggest that teleoperation data can provide supervision for converting action-free kinematic demonstrations into useful training data for robot policies. More videos are available on https://generativeactionsynthesis.github.io.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.