acceptodds
Under review as a conference paper at ICLR 2027

A Perfect Model Is Not Enough: A case study on long horizon planning in LeWorldModel

Abstract

We study what limits long-horizon planning in JEPA world models through a case study on LeWorldModel (LeWM), a compact world model trained end-to-end offline on pixels whose evaluation protocol has become standard in subsequent work. Several of LeWM’s successors attribute its failures at long horizons to compounding rollout error. We first show that extensions of the standard protocol that simply increase the goal offset cannot test this claim: seemingly minor details of several environments and datasets make them poor tests of long-horizon planning at larger goal offsets. On environments that avoid these pitfalls, we evaluate LeWM on held out episodes across increasing start-goal offsets, varying its rollout source and terminal cost separately. Our results show that (a) planning with a perfect model (by executing each candidate in the simulator) improves long horizon success only modestly, while (b) replacing only the Euclidean terminal cost with a cost-to-go learned on frozen embeddings can achieve similar or greater outcomes without retraining the world model. These insights suggest that evaluation protocol and planner costs inherited by follow up work deserve as much scrutiny as the predictive world model itself.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.