ReWorldModel: Rethinking Visual Representation Learning in Latent World Models
Abstract
Latent world models learn compact visual representations for predicting action-conditioned future states. However, visual information is organized hierarchically across encoder depth: low-level features retain fine-grained spatial and geometric details, favoring representational completeness, whereas high-level features form compact abstractions that favor predictability. Existing world models typically operate only on final-depth representations, potentially discarding details required for precise control. This creates a cross-level trade-off between representational completeness and predictability. To address this, we introduce ReWorldModel (ReWM), a multi-level representation architecture centered on Mixture-of-Depths Attention (MoDA). MoDA enables high-level queries to selectively retrieve low-level features from earlier encoder layers, while a modernized Transformer design improves optimization and spatial modeling. Full ReWM achieves 100% success on four of five tasks and 99.0% mean success on PushT, reaching 95% on PushT with 5× fewer processed training sequences than LeWM under its official recipe. Under a shared training recipe with three paired seeds, a controlled ReWM variant improves Reacher success by 74.4 percentage points at eight encoder updates and reaches 95% with 8× fewer updates on Reacher and TwoRoom. Layer-wise probing further shows that ReWM progressively organizes control-relevant position, orientation, and interaction information across encoder depth, producing final predictive representations with richer control-relevant structure. ReWM thus reshapes the completeness–predictability trade-off toward representations that maintain strong latent predictability while becoming useful for downstream control substantially earlier. These results demonstrate that cross-depth access improves the completeness–predictability frontier of latent world-model representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.