Beyond Single Goal Images: Learning Goal Metrics for LeWorldModel Planning
Abstract
A single goal image shows both task-relevant information and incidental visual details, although many visually different outcomes may solve the same task. We extend LeWorldModel (LeWM) to plan toward sets of successful outcomes by separating predictive representations from goal-specific cost evaluation. We freeze the encoder and action-conditioned latent dynamics and learn a goal projector, used as either a linear map or a nonlinear network. The distance between a projected latent vector and the mean of projected goal examples defines the planning cost. Training combines compactness of successful examples with distribution regularization, temporal ordering along goal-reaching trajectories, and preferences for demonstrated action sequences over corrupted alternatives. The resulting cost integrates naturally into the existing latent model-predictive-control procedure without modifying the world model. Across three world-model seeds, the nonlinear projector achieves mean success rates of %, %, and % on Push-T, Two-Rooms, and OGBench-Cube, respectively. It surpasses both single-goal-image LeWorldModel planning and latent 1-nearest-neighbor goal-set planning, improving over the latter by , , and percentage points under the same evaluation protocol. These results support learning task-specific planning costs while preserving a shared predictive representation. Source code will be made available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.