acceptodds
Under review as a conference paper at ICLR 2027

TacWM: Unified World Modeling of Vision, Touch, Action, and Object Pose for Dexterous Manipulation

Abstract

High-performance dexterous manipulation requires accurate modeling of the contact dynamics between the hand and manipulated objects. However, existing methods either overlook the importance of tactile feedback in contact-rich tasks or focus solely on fitting action trajectories without modeling object status, thereby limiting manipulation performance. We introduce TacWM, a unified multimodal world model that jointly models visual observations, tactile feedback, actions, and object poses. Rather than modeling each modality independently, TacWM utilizes a unified diffusion transformer to model their joint distribution under causal temporal conditioning. Specifically, each modality is first encoded by a dedicated encoder, and TacWM is trained to predict the next-step representations of all modalities from multimodal history. This joint prediction allows different modalities to complement and reinforce one another. In particular, predicting future tactile feedback and object poses helps the model anticipate contact changes and object motion, thereby improving action generation for dexterous manipulation. To support large-scale training on heterogeneous datasets, we also propose a cross-embodiment mechanism that progressively aligns action and tactile signals from new embodiments in shared latent spaces. Extensive experiments demonstrate that TacWM achieves state-of-the-art performance in dexterous manipulation, marking a step toward general-purpose multimodal dexterous manipulation. Beyond enabling effective dexterous manipulation, TacWM can also effectively predict future videos, tactile feedback, and object poses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.