DasWM: Disentangling Dynamic and Static Factors for Latent World Modeling
Abstract
We introduce DasWM, a world model representation that softly disentangles dynamic, action-relevant factors from static ones, allowing the allocation of prediction capacity to the dynamic parts. Latent world models make visual planning efficient by encoding observations (e.g. pixels) into compact latent representations. However, not all encoded information needs to be subject to prediction. We show in this work that recently proposed JEPA-like latent world models fail when static visual factors are introduced. We identify the cause of this failure as static-dynamic entanglement in the representation. Our DasWM latent world model encodes each observation as a sequence of tokens, ordered with respect to their dynamic responsiveness. This leads to a representation where dynamic information lies in the early tokens, and static information in the latter ones, minimizing the entanglement. We evaluate DasWM on background-modified versions of several datasets and demonstrate clear improvements in planning, in addition to other benefits such as flexible budget planning at test-time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.