Latent World Models Need to Also See the Details: Keyframe-Aware Visual Alignment
Abstract
Latent world models built on the Joint Embedding Predictive Architecture (JEPA), such as LeWorldModel, learn environment dynamics entirely in latent space and avoid the cost of pixel-level reconstruction. This purely latent paradigm, however, has two fundamental limitations. In fine-grained manipulation, the target object typically occupies only a tiny fraction of the frame, and the encoder's information bottleneck can suppress visual details critical to success. Moreover, the absence of any decoder that maps latents back to an observable space makes the long-horizon accuracy of latent rollouts unverifiable, allowing an overconfident yet physically implausible plan to pass unnoticed. We introduce **KAVA** (Keyframe-Aware Visual Alignment), a hybrid framework that augments a JEPA next-embedding objective with a temporal-saliency key-frame selector—combining the JEPA-loss gradient norm with an action-perturbation uncertainty score—and applies a pixel-level reconstruction loss only to a small subset of information-dense frames per trajectory. The reconstruction gradient is stopped at the predictor, preserving the latent signal the JEPA objective is designed to learn, and the decoder simultaneously serves as an online probe for the consistency of imagined rollouts. Across four image-based control benchmarks, KAVA achieves the best planning success rate among four representative latent world-model baselines. Selective, saliency-driven observation supervision is therefore a promising recipe for retaining the efficiency of reconstruction-free world models without sacrificing visual fidelity or rollout verifiability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.