Learning Physically Grounded Latent Actions from Human and Robot Videos
Abstract
Latent-action pretraining provides a scalable way to learn robot policies from actionless videos, but existing objectives are driven primarily by visual change and therefore provide only indirect supervision of physical interaction. Tactile and force measurements can provide direct physical signals, but they are available for only a limited subset of existing human and robot videos. Our key insight is to decouple the acquisition of physical supervision from its transfer to latent-action learning: physical supervision derived from a limited corpus of instrumented interactions can be propagated to broader collections of human and robot videos without paired sensor measurements. Building on this insight, we propose PHYSLA, a Physically grounded Latent-Action framework that learns a shared, horizon-structured representation from heterogeneous physical measurements, including pressure distributions, tactile marker displacements, and joint torques, and infers these representations from causal RGB observations. During vision-language-action pretraining, the inferred physical targets and future visual prediction jointly supervise the same latent actions. Importantly, the transferred targets capture interaction-relevant latent structure rather than calibrated sensor readings, allowing physical supervision to scale beyond the instrumented corpus while requiring no tactile or torque sensing at deployment. PHYSLA achieves 98.2% average success on LIBERO, 74.0% and 64.7% on SimplerEnv’s Google Robot and WidowX settings, respectively, and 65.5% across five real-world G1 tasks. These results support the effectiveness of transferring measurement-derived physical supervision to latent-action learning from human and robot videos
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.