Video Generators Can Predict Touch: Tactile Synthesis from Egocentric Video
Abstract
Tactile sensing matters for dexterous manipulation. However, tactile data is expensive and slow to collect at scale. In this work, we investigate how the interaction priors learned by video generators can be transferred to touch prediction from egocentric videos. We find that Wan2.1-Fun-Control can be easily repurposed for tactile pressure prediction via lightweight LoRA fine-tuning with its base weights frozen. On the OpenTouch test set, our model, using only video observations, achieves an MSE of , hard IoU of , and soft IoU of for tactile prediction. This setup also extends naturally to joint prediction of tactile pressure, hand configuration, and hand–object segmentation. We additionally evaluate tactile prediction on the bimanual EgoTouch dataset and qualitatively assess transfer to in-the-wild video clips without further fine-tuning. A randomly initialized generator fails to recover tactile structure in our control experiment, supporting the benefit of generative pretraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.