Learning Visuomotor Policies from JEPA Latent Transition Models
Abstract
Video pretraining offers more than visual representations: predictive models also learn how visual states evolve. The success of recent generative world-action models (WAMs) suggests that pretrained transition models can themselves provide a transferable foundation for policy learning. We ask whether this principle extends to latent transition models learned through joint-embedding predictive architectures (JEPAs): beyond transferring the visual encoder, can a predictor that consumes robot actions be repurposed into a visuomotor policy that produces them? We introduce LTM-Policy, which adapts the latent transition model through two complementary routes. A policy route employs learnable action queries to generate robot actions, while a world route retains action-conditioned prediction of future representations from demonstrated transitions. At deployment, the resulting policy generates actions without pixel reconstruction, future-video rollout, or model-predictive search. Controlled experiments show that pretrained transition priors remain beneficial after repurposing the model for action generation. Retaining future-latent prediction during adaptation provides an additional 5.4-point benefit, and the transfer advantage persists in low-data regimes with limited robot demonstrations. Across the four LIBERO suites, LTM-Policy achieves 97.4% average success using only a single forward pass at deployment, without pixel reconstruction or future-video rollout. We further evaluate the method on LIBERO-Plus and real-world manipulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.