acceptodds
Under review as a conference paper at ICLR 2027

Time-JEPA: Grounding Video World Models with Physical Time Series

Abstract

Video world models learn how visual environments evolve by predicting miss- ing or future content. However, physical-state information can provide useful context for predicting how environments evolve. Prior work incorporates this in- formation through different mechanisms: domain-specific physical laws can com- plicate multi-domain training, while supervision on individual frames or frame pairs gives limited attention to how physical state evolves over time. We intro- duce Time-JEPA, which fine-tunes a video world model using low-dimensional, temporally aligned physical time series as auxiliary supervision. Physical time series connect video representations with physical state at corresponding tempo- ral positions, encouraging the model to preserve how physical state evolves. In multiple domains including autonomous driving, robot manipulation, and human motion, Time-JEPA improves current-state decoding, multi-horizon future-state prediction, and kinematic reconstruction over matched video-fine-tuned controls across both condition and state-variable distribution shifts. Multi-domain training provides aggregate gains over matched single-domain training in all three tasks. External evaluations further suggest that time-series supervision improves sensi- tivity to physical anomalies and the representation of physical interactions, includ- ing in unseen scenarios. Aggregate gains on a second video backbone indicate that the method can be applied to more backbones. These results suggest that physical time series help video world models better understand the physical world.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.