Sample Efficient World Models using Gradient Penalized Latent Dynamics
Abstract
World models are trained on one-step transitions and then rolled out for dozens of steps to train a policy. Where those transitions are sparse, the learned transition function can change rapidly between neighboring states, and long imagined rollouts compound the resulting errors. In continuous control, neighboring states evolve under similar transition laws, so a transition observed at one state is informative about its neighbors. To exploit all such information, we propose GPLD (Gradient Penalized Latent Dynamics), a method that makes possible the sharing of this information between neighboring states. We derive it as the small-neighborhood limit of a finite-difference penalty on transition laws in a finite MDP, which yields the squared Frobenius norm of the dynamics Jacobian. We apply it to DreamerV3’s posterior latent distribution, the map through which observations shape the learned dynamics, and estimate it with Hutchinson-style stochastic probes. With a single configuration held fixed across every task, GPLD raises final return over DreamerV3 on 15 of 18 proprioceptive DeepMind Control tasks (sign test p = 0.0075) and mean final return by 9.8% on the 12 tasks where DreamerV3 has not already saturated (18.1% in normalized return), and by 20.7% on the six humanoid and dog tasks. On quadruped it reaches high return earlier, and its multi-step prediction error is lower than DreamerV3’s on most measured quantities and significantly higher on none. Sharing information between neighboring states, as proposed here, improves the sample efficiency of world models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.