acceptodds
Under review as a conference paper at ICLR 2027

ACTS: Learning Visuo-Tactile World Models from Human Interaction

Abstract

Action-conditioned world models for contact-rich manipulation require data that captures how contact forms, slides, and releases under varying force. Such data are slow to collect through robot teleoperation, which often lacks force feedback and is constrained by gripper geometry. We ask whether these dynamics can instead be learned from human physical interaction and transferred to a robot. We present ACTS, a visuo-tactile world model trained only on human interaction data. Operators manipulate handheld GelSight sensors while we record third-person and wrist video, tactile images, 6-DoF sensor motion, and tactile-derived normal force, yielding React, a 4.9-hour bimanual dataset across three objects. We represent actions at the tactile sensor as motion augmented by force change, providing a shared action space for human collection and robot execution. ACTS jointly predicts future scene, wrist, and tactile observations. On held-out human data, ACTS improves scene-view prediction by 0.7–2.0 dB PSNR over two visuo-tactile world models trained from scratch on React. Without fine-tuning, predicted tactile imprints on robot teleoperation still appear and disappear with recorded contact. Visual transfer is less robust: camera predictions match or exceed frame copying on two of three objects, while wrist views can blur or hallucinate structure. These results suggest that contact dynamics transfer more reliably than visual appearance from human to robot. Project website and supplementary videos: https://acts-anon.github.io/

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.