Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
Abstract
World generative models are typically used through their outputs: a rendered future, a video-conditioned action, or latent context from a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. As a generator transforms a corrupted future into a coherent trajectory, its intermediate states organize appearance, spatial layout, and interaction across abstraction levels. Can this computation be internalized in a representation inferred from the present alone? We present Enfold, which transfers it into a representation predicted from the current visual context and language instruction. During training, multi-level states exposed while the generator processes an observed future supervise a current-only encoder. The representation conditions future generation and is read by task heads without allowing task gradients to reshape the encoder. At deployment, action prediction bypasses the generator. This asymmetry encourages the encoder to retain transition structure predictable from the present rather than stochastic, future-dependent nuisance variation. Across LIBERO, RoboTwin2.0, and real-robot tasks, Enfold supports strong control and cuts action latency by relative to Fast–WAM; Enfold-Flash reaches . Analyses show that it suppresses nuisance variation and preferentially captures changes emerging over longer horizons. When a human alters the current scene, both the generated continuation and executed actions adapt, inconsistent with fixed trajectory replay. These results recast a world generator as a source of predictive control representations: its future need not be materialized at every step if its internal structure can be enfolded into the present.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.