acceptodds
Under review as a conference paper at ICLR 2027

WorldPrior: Internalizing Multimodal Semantics into Latent World Models

Abstract

Latent world models learn compact predictive representations of visual dynamics, but matching future features alone does not guarantee that their hidden states expose the object identity, affordance, interaction phase, relational structure, and contextual cues that determine which futures matter for embodied prediction and control. Vision-language models (VLMs) contain complementary multimodal priors, yet using them online changes the deployment interface and introduces a large inference-time cost. We introduce WorldPrior, a training-only adaptation framework that internalizes multimodal semantics into a latent world model while preserving a VLM-free deployment graph. A frozen VLM processes only the observed clip, and detached observation-side representations are aligned with selected internal predictor states through cosine distillation. Because high-level multimodal supervision can perturb fine spatial and dynamical structure, a frozen reference predictor initialized from the same baseline checkpoint acts as a functional anchor, while the backbone-native predictive objective remains active throughout adaptation. All VLM-specific projections, semantic readouts, reference branches, and training-only adaptation objectives are removed at inference. On EgoDex, WorldPrior reduces ADE/FDE from 0.0705/0.0646 to 0.0597/0.0572 and raises threshold accuracy from 0.436 to 0.564. On EgoPAT3D, it obtains the best or tied-best result on all four 3D metrics and three of four 2D metrics across seen and unseen scenes. With LeWM as the control backbone, average success increases from 85.5% to 90.0%; the same principle transfers to DINO-WM, improving the matched four-task average from 83.0% to 85.5%. WorldPrior retains the inference graph and practical latency of the corresponding VLM-free backbone, whereas our online VLM-guidance reproduction is more than 22× slower in forecasting. These results support a broader view of multimodal foundation models as training-time representational priors that can strengthen perception, relational understanding, and task-relevant inference in latent world models without becoming deployment-time dependencies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.