Learning Gravity from Video: Structured Latent States for Prediction and Intervention
Abstract
A physical system evolves according to its current state and governing dynamics. This suggests organizing video world models around visually inferred states and reusable evolution rules. Yet predictive learning alone does not ensure that learned representations support sustained physical evolution and meaningful intervention. We introduce GravJEPA, a self-supervised world model built on LeWorldModel that separates visual state estimation from dynamical evolution. The model infers structured latent components from video histories and propagates them with a time-homogeneous Koopman-form transition. This formulation supports direct queries at future times after observations stop. To align the inferred states with this dynamical structure, we introduce regularization objectives that encourage translation equivariance and photometric invariance in position-associated representations. The visual states are learned without ground-truth physical-state or gravity supervision. On controlled gravity videos, GravJEPA achieves lower trajectory prediction errors than an autoregressive baseline trained with the same equivariance and invariance regularization. It also outperforms this baseline in long-horizon prediction and out-of-distribution tests with unseen velocity and gravity magnitudes. Beyond prediction, matched interventions on inferred velocity- and gravity-associated components produce corresponding changes in future motion through the same transition. Together, these results establish a shared latent-state interface for forecasting and intervention, taking a step toward physics-engine-like world models learned from video.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.