Beyond Terminal Matching: Arrival-Aware Planning with World Models
Abstract
Planning with a learned world model involves deciding not only how to reach a goal, but also when to reach it. Matching the goal only at a fixed terminal step does not explicitly favor earlier arrival. We introduce an arrival-aware planning objective for learned world models that requires neither model fine-tuning nor training of auxiliary components. The objective treats each prediction step as a candidate arrival time and assigns it a cost combining elapsed time with a penalty for predicted goal mismatch. These costs are aggregated through a smooth minimum, allowing action optimization to favor earlier goal attainment without prescribing a fixed arrival step. We analyze how this objective adaptively weights gradient contributions from different candidate arrival times. Experiments on visual goal-reaching tasks show reductions in budgeted physical interaction steps under a shared planning configuration. Controlled comparisons separate flexible endpoints, which improve across four LeWM tasks, from explicit time preference, which provides further gains on TwoRoom and Reacher. These findings highlight arrival-aware objectives as a means of improving physical-interaction efficiency without additional training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.