acceptodds
Under review as a conference paper at ICLR 2027

VPFlow: Value-Predictive Planning with Imperfect World Models for Offline Goal Reaching

Abstract

Offline goal-conditioned reinforcement learning turns unlabeled trajectories into reward-free skill acquisition: relabel future observations as goals, and one dataset supervises many tasks. Model-based planning could add test-time compute on top, but existing planners share an assumption: the world model must predict the next state, in pixels or in latents, faithfully enough to plan with. We challenge this assumption. Offline goal reaching admits a closed-form Monte Carlo return under sparse rewards, so value learning becomes regression instead of temporal-difference bootstrapping; we use this return as the sole supervision for the latent dynamics, training imagined latents to predict goal-reaching value rather than the next state, which prevents representation collapse without stop-gradients or target encoders. We present VPFlow, which pairs this deliberately imperfect world model with a flow policy and steers it at test time in the policy's noise space, avoiding the out-of-distribution action queries of action-space planning. On 30 OGBench datasets with state or pixel inputs, VPFlow raises mean success from to (state) and from to (pixel) over the strongest baselines, even though its state-prediction error is orders of magnitude larger and its imagined latents stray far from the manifold of encoded states; reducing that error lowers success. Planning needs value-predictive representations, not faithful simulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.