LaFree World Model: Label-Free Pretraining for Action-Conditioned JEPA World Models
Abstract
Learning world models from observation sequences could reduce dependence on synchronized action labels, but learning to explain observed transitions does not necessarily yield dynamics that support downstream control. We introduce LaFreeWM, an end-to-end action-conditioned JEPA world model that learns dynamics using discrete latent actions inferred from consecutive observations. A quantized action interface, together with latent transition objectives, couples action discovery and representation learning without action annotations. A small action-labeled subset subsequently adapts the pretrained world model to the environment's control inputs for planning. Across control benchmarks, LaFreeWM improves planning with limited action supervision while retaining a compact efficient planner. Ablations show the benefit of end to end training and reveal a prediction–planning mismatch: removing quantization substantially lowers held-out latent prediction error while degrading downstream planning. The results demonstrate the value of label-free pretraining for counterfactual planning with action-conditioned world models and highlight the importance of evaluating inferred-action representations through the control they enable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.