Horizon-Aligned Scoring: A Longer View for Short-Horizon World Model Planning
Abstract
LeWorldModel (LeWM) learns visual latent dynamics from reward-free trajectories and uses the cross-entropy method (CEM) to plan toward image goals. It ranks candidate action sequences by their predicted endpoint-to-goal distances, so short rollouts can undervalue preparatory actions whose benefits emerge later. We introduce Horizon-Aligned Scoring (HAS), which uses an action-free latent predictor to extend each endpoint toward the goal's time scale before scoring. Trained on LeWM's existing trajectories, the predictor adapts its recursion depth to the remaining temporal gap, given the goal offset. HAS thereby extends candidate evaluation without lengthening the action-conditioned search. Across PushT, TwoRoom, and Reacher, HAS improves mean success when the goal image is sampled beyond LeWM's default planning span of 25 environment steps. For goal images sampled 50 environment steps after the initial observation, PushT success rises from 42.8% to 67.8%; at a 100-step goal offset in TwoRoom, HAS reaches 70.4%, compared with 55.4% for the matched hybrid variant of Trajectory Reachability Metrics (TRM-Hybrid) and 16.4% for LeWM. The gains persist across three independently trained checkpoints per task. HAS also outperforms a LeWM control with twice the action horizon at every tested longer goal offset under the same replanning interval. Together, the results show that HAS improves LeWM's long-range planning without additional training data or a longer action-conditioned planning horizon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.