acceptodds
Under review as a conference paper at ICLR 2027

RoboTryOn: Synthesizing Dexterous Robot Demonstrations from Human Videos via Virtual Try-On

Abstract

Robot foundation models have shown remarkable performance in robot control, yet adapting them to a new task still relies heavily on teleoperated demonstrations. Egocentric human videos offer a scalable substitute, but retargeting human poses into robot actions or overlaying rendered robots on human leaves significant visual or action gaps for dexterous hands with high-DoF. To address this limitation, we propose RoboTryOn, which treats human-to-robot translation as virtual try-on: try-on keeps the person and swaps the garment, RoboTryOn keeps the demonstrated motion and swaps the human for the robot. Specifically, RoboTryOn synthesizes robot data using masks and hand skeletons as conditioning signals, while reference images specify the target robot embodiment. To maintain visual consistency, we translate selected keyframes autoregressively and use them to guide video synthesis conditioned on the skeleton. Additionally, we use an inverse dynamics model (IDM) that incorporates depth features to infer action labels from synthesized videos for policy learning. Our experiments across multiple humanoid embodiments show that policies trained with our synthesized data improve the average success rate on novel tasks, including challenging bimanual manipulation, from 2% to 60%, without any real robot demonstration of those tasks. Moreover, training with our synthetic data matches the performance of policies trained solely on teleoperated demonstrations at approximately a third of the collection cost, demonstrating the cost efficiency of our approach.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.