When Does Temporal Distance Improve World-Model Planning?
Abstract
Latent world models plan by keeping candidate action sequences whose predicted final state lies closest to the goal in latent space. This latent distance, however, is not trained to reflect reachability: states that are close in latent space may still require many steps to connect. Learned temporal distances can address this mismatch, but it remains unclear when replacing the latent distance will improve planning, or whether the benefit can be predicted before planning. We propose a simple offline diagnostic using trajectories already collected to train the world model. A useful planning distance should rank states requiring fewer steps to reach the goal as closer; we therefore measure how much a learned temporal distance improves this ordering over the original latent distance. With the world model and search fixed and the execution schedule matched, planning improves substantially when the temporal distance greatly improves temporal ordering, while gains are smaller when the original latent distance already captures that ordering. We further apply the diagnostic before planning in held-out settings, where its predictions match the observed distinction between large and smaller planning gains. Controlled experiments show that the large gains come from learning temporal structure rather than simply changing the representation, search budget, or execution schedule. The results suggest that temporal ordering in offline trajectories provides a practical way to identify when latent distance is a bottleneck in world-model planning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.