DO LATENT WORLD MODELS REALLY SELECT THE RIGHT ACTIONS?
Abstract
Latent world models have recently shown strong results in planning from images, particularly on navigation and manipulation tasks. They score candidate action sequences by the distance between their predicted embeddings and the embedding of a goal image. We find that on fixed candidate sets this score usually misses the best action sequence, and that on the manipulation tasks most models do no better than a random choice. On Push-T the released model's predictor causes this. If these scores misrank fixed candidates, why does planning still work so well? Part of the answer is that the same scores rank the planner's own candidates far above a random choice, with the best one on top in 18.1% of Push-T searches against 1.6%, and we show that wider search and more frequent replanning also contribute to this success. PointMaze success falls from 91% to 76% without replanning. On Push-T, a wider search raises success without replanning from 13% to 67%. Further analysis on PointMaze finds a successful action sequence among the planner's candidates in over 90% of episodes, yet the candidate the model scores highest succeeds in 63%. Our findings suggest that strong results in planning do not show that a latent world model selects the right actions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.