Frozen in a Frame: The Velocity Blind Spot in JEPA World Models
Abstract
Joint-embedding predictive architectures for world modeling train an encoder so that a predictor can map a current frame embedding and an action to the embedding of the next frame. The target for this prediction always comes from a single rendered frame. We show that this choice has a structural blind spot. Every renderer without motion blur is instantaneous: it draws a scene from its configuration alone and leaves no trace of how fast that configuration is changing. So a single-frame embedding is a deterministic function of configuration alone, and it carries no information about velocity. This holds for the optimal predictor under any encoder, however large, including the official released LeWM weights. We confirm this on official checkpoints across four real benchmarks, PushT, Reacher, Cube, and TwoRoom, where every linear velocity probe sits at or below its own chance floor while position probes reach . We then introduce RateIdent, a three-stage diagnostic protocol, and TI-JEPA, a lightweight fix that splits the latent into a pose code and an explicit finite-difference motion code and predicts both. Across three physically grounded environments, TI-JEPA gives a statistically significant, seed-robust improvement on a stop-at-goal planning task over a matched-memory baseline, for example a 55% lower final distance on Pendulum () and 64% lower on CartPole (). We reproduce the theory's central prediction at official ViT-Tiny plus AdaLN-transformer scale on two environments, then push the same recipe onto real dm_control Reacher photographs trained entirely from scratch, where TI-JEPA's kill-experiment branch separation exceeds the memory-having baseline's by roughly , the largest margin in the paper. We also compare directly against a recurrent RSSM-style predictor at the same footprint. Its opaque hidden state matches or beats TI-JEPA's raw rollout accuracy on two of three environments, but it cannot be probed for pose and motion separately the way TI-JEPA's explicit split can, and it loses outright on the environment with the most coupled dynamics. Together, a checkable formal argument and six evaluated environments show that single-frame targets are the wrong object to predict when velocity matters, and that a small, interpretable structural change fixes it without any privileged supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.