acceptodds
Under review as a conference paper at ICLR 2027

ME-Dex 1.0: Unifying Heterogeneous Tactile Sensing into World Action Modeling

Abstract

World Action Models predict future visual states to inform robot action generation. Tactile sensing directly captures contact states and force changes during physical interaction. Existing methods often treat tactile sensing as a conditioning input for action generation, thereby modeling only the current contact state. Instead, tactile signals should be modeled alongside video as future observations of the evolving world state. We present ME-Dex 1.0 (ME-Dex), a unified World Action Tactile Model for joint modeling of visual states, tactile states, and robot actions. ME-Dex adopts a Mixture-of-Transformers comprising a Video Expert, a Tactile Expert, and an Action Expert, all trained with flow matching. Multi-modal joint attention connects the three experts at intermediate layers, allowing action generation to exploit learned visual and tactile dynamics during joint generation. To support heterogeneous tactile inputs across embodiments, ME-Dex maps sensor surfaces to a Canonical Hand Model and encodes the aligned observations into a unified latent space with a Unified Tactile Autoencoder. We augment tactile-free simulation trajectories with synchronized tactile observations for joint visual, tactile, and action learning. Experiments on RoboTwin, DexJoCo, and ManiFeel improve average manipulation success over the strongest baselines by 11.2, 8.9, and 15.0 percentage points. Qualitative rollouts on physical robots further demonstrate the use of visual and tactile observations across grippers and dexterous hands. Code and supplementary materials are available at https://anonymousmedex.github.io/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.