acceptodds
Under review as a conference paper at ICLR 2027

PreShort: Efficient World Modeling with Pretrained Visual Representations

Abstract

Model-based reinforcement learning (MBRL) leverages learned world models for sample-efficient control, provided that their representations retain the information needed to predict the consequences of actions. General-purpose pretrained visual representations offer reusable features learned beyond an agent's online experience, motivating their use as a prediction space for world modeling, while turning these features into imagined trajectories requires learning action-conditioned transitions and accounting for uncertain future states. We introduce PreShort, an online world model that employs frozen visual features as both observation representations and continuous imagination targets. Its core component, Differential Shortcut Dynamics, encodes latent-action history with differential attention and conditions a stochastic shortcut predictor on the resulting context. The context is reused across a few refinement steps to generate each successor, and imagined trajectories support actor-critic learning. Freezing the visual interface also enables feature caching and confines online world model updates to downstream modules. Experiments show that PreShort achieves a mean human-normalized score of % on Atari K, a mean return of on the DeepMind Control Suite, and a return of on Crafter, exceeding strong world model baselines in mean aggregate performance across all three benchmarks. Further analyses assess rollout accuracy, compatibility with different pretrained visual backbones, and computational efficiency in terms of trainable parameters and training cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.