acceptodds
Under review as a conference paper at ICLR 2027

Horizon-Aligned Scoring: A Longer View for Short-Horizon World Model Planning

Abstract

LeWorldModel (LeWM) learns visual latent dynamics from reward-free trajectories and uses the cross-entropy method (CEM) to plan toward image goals. It ranks candidate action sequences by their predicted endpoint-to-goal distances, so short rollouts can undervalue preparatory actions whose benefits emerge later. We introduce Horizon-Aligned Scoring (HAS), which uses an action-free latent predictor to extend each endpoint toward the goal's time scale before scoring. Trained on LeWM's existing trajectories, the predictor adapts its recursion depth to the remaining temporal gap, given the goal offset. HAS thereby extends candidate evaluation without lengthening the action-conditioned search. Across PushT, TwoRoom, and Reacher, HAS improves mean success when the goal image is sampled beyond LeWM's default planning span of 25 environment steps. For goal images sampled 50 environment steps after the initial observation, PushT success rises from 42.8% to 67.8%; at a 100-step goal offset in TwoRoom, HAS reaches 70.4%, compared with 55.4% for the matched hybrid variant of Trajectory Reachability Metrics (TRM-Hybrid) and 16.4% for LeWM. The gains persist across three independently trained checkpoints per task. HAS also outperforms a LeWM control with twice the action horizon at every tested longer goal offset under the same replanning interval. Together, the results show that HAS improves LeWM's long-range planning without additional training data or a longer action-conditioned planning horizon.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.